Published on 18 September 2026
Does an AI clinical decision support tool reduce treatment failures in primary care? A Kenyan cluster-randomized trial finds no difference
The essentials
Sixteen primary care facilities in Kenya, 103 randomized clinicians, 9,347 consultations analyzed: a generative AI decision support system built into the electronic medical record was tested against usual care. The primary endpoint was not a quality score but a composite of treatment failure at 14 days, adjudicated blind by an independent expert panel. It occurred in 102 of 4,693 patients (2.2%) in the assisted arm and 94 of 4,654 (2.0%) in the control arm, an adjusted odds ratio of 0.77 (95% CI 0.55 to 1.08; p = 0.13). The trial does not demonstrate a reduction in treatment failure. It does report a clear improvement in clinical documentation quality, a small drop in antibiotic spending, and no safety signal. It carries targeted conflicts of interest, it took place in a private network of urban clinics that is not our health system, and it has nothing to do with psychiatry. Its value for a French-speaking reader is methodological, not clinical.
The context
Claims made for generative-AI decision support have long outrun their evaluation in real practice. The most cited demonstrations rely on clinical vignettes and in silico methods, that is, on setups where the model is judged outside the actual flow of care. The authors state their starting point cautiously, and that caution is worth quoting exactly: rigorous evidence on the performance of large language models “in real-world, low-resource clinical settings remains limited.” They do not say such evidence is absent, and they explicitly place their work in the continuation of recent randomized evaluations conducted in other clinical settings.
Psychiatry is not the field this trial tested, and that is exactly why it belongs here. The question it asks, whether a decision support tool actually changes what a hurried clinician does, will arise in the same terms once these tools reach psychiatric consultations.
The study at a glance
| Population | 9,347 consultations, all ages, 16 primary care facilities |
| Penda Health private network, Nairobi and Kiambu counties, Kenya. Despite the trial being framed around adults, 3,553 of 9,347 patients (38%) were under 18; 5,658 (61%) were 18 to 55 years old and 136 (1.5%) were over 55. Presenting complaints were varied, dominated by febrile or infectious illness (5,585; 60%). Neurological and psychiatric diagnoses accounted for 1,242 patients (13%). | |
| Intervention | “AI Consult” version 2.0, built into the electronic medical record |
| Developed by Penda Health. The underlying model is OpenAI’s GPT-4o (May 2025 version), temperature 0.1, top-p 1.0, output capped at 1,024 tokens. The system analyzes data entered during the visit and generates recommendations without the clinician issuing a query, with a three-tier visual signal: green, amber, red. | |
| Comparator | Same facilities, same electronic medical record, feature disabled |
| “AI Consult” was switched off for control-arm clinicians. Usual care, with no additional incentive to follow recommendations. | |
| Primary endpoint | Treatment failure within 14 days, expert-adjudicated composite |
| A repeat visit for unresolved symptoms, unplanned escalation to a higher level of care or to emergency services, or a safety event (delayed or missed referral, inappropriate prescription, missed diagnosis, life-threatening event, death). | |
| Secondary endpoints | Documentation quality, sentinel conditions, prescribing, satisfaction |
| Clinical documentation quality, management of sentinel conditions (hypertension, type 2 diabetes, malnutrition), appropriateness of antibiotic and antimalarial prescribing, patient satisfaction. Consultation duration and per-category drug costs were post hoc analyses, not prespecified secondary endpoints. | |
| Design | Pragmatic, multicenter, parallel-group cluster-randomized trial |
| The unit of randomization was the clinician (“clinical officer”), not the facility: 103 clinicians randomized, 52 to the assisted arm and 51 to the control arm, patients nested within clinicians and facilities. Enrollment ran from April 22 to July 16, 2025. Protocol and statistical analysis plan published in advance. Oxford level of evidence 1b. | |
Quality control
| Registration and protocol | Compliant and documented |
| Trial registered with the Pan-African Clinical Trials Registry under number 202502499779176. Protocol and statistical analysis plan deposited publicly on Zenodo. Reporting follows the CONSORT-AI and CONSORT 2010 extensions for cluster-randomized trials, with a completed CONSORT checklist in the supplementary material. | |
| Randomization unit | Appropriate, 103 clusters |
| Randomization at the clinician level is the coherent choice for a tool whose use creates an individual learning effect: a single practitioner cannot alternate between arms without contamination. Variable block randomization (blocks of 4, 6, or 8) performed by the trial statistician. A cluster count of 103 is comfortable. The limitation lies elsewhere: both arms coexisted within the same facilities, and the authors acknowledge that informal exchange of clinical reasoning between colleagues could not be entirely prevented. | |
| Sample size and power | Target reached, power limited |
| The target of 9,000 consultations was calculated to detect a 50% relative reduction in treatment failure, from 2% to 1%, with a design effect of 1.5, 80% power, a two-sided alpha of 0.05, and 10% loss to follow-up. That target was reached. But the authors write that the observed event rate was lower than expected and that the trial was powered for an effect larger than the one observed, hence limited precision. Post hoc simulations put the sample size needed to detect modest differences in rare clinical events at over 100,000 patients. | |
| Nature and collection of outcome data | Independent adjudication |
| The primary endpoint is not an administrative record. It rests on blinded telephone follow-up at day 3 and day 14 by research assistants, then adjudication by an independent panel of six Kenyan family physicians with 10 to 16 years of practice. Each event is reviewed independently by two panel members, with a third resolving persistent disagreement. Adjudicators had no access to treatment allocation or to the model’s outputs. | |
| Blinding | Partial, properly organized |
| Blinding of clinicians was impossible, an inherent constraint of this kind of intervention and not a conduct flaw. Patients, the research assistants conducting follow-up, the adjudication panel, and, per the reporting summary, the analysts for the primary analyses, were all blinded to allocation. | |
| Funding and conflicts of interest | Targeted, disclosed, worth noting |
| Funded by the Gates Foundation, grant INV-068056, sponsored by PATH, Seattle. The authors state the funder had no role in design, data collection, analysis, the decision to publish, or manuscript preparation. Two authors, R. Korom and S. Kiptinness, hold stock options in Penda Health, the organization where the trial was conducted. OpenAI, which owns the model, provided in-kind support to Penda Health, compute credits and technical guidance on its API, for developing and tuning “AI Consult”; the authors state the decision to use OpenAI’s product predated that offer and that OpenAI had no role in the study or the publication decision. The remaining authors declare no competing interests. | |
| External validity | Limited, acknowledged by the authors |
| The authors themselves write that conducting the trial in a single private network of urban Nairobi clinics may limit generalization to rural, peri-urban, or public-sector settings. They add that the network’s already high baseline quality of care, maintained through regular clinical audits and peer review, likely reduced the measurable room for improvement. No direct extrapolation to French psychiatric practice is possible. | |
The results
Adjusted odds ratio for treatment failure at 14 days, 95% CI 0.55 to 1.08, p = 0.13. The point estimate favors the intervention, but the interval crosses the null value: no reduction is demonstrated.
| Treatment failure at 14 days | 102/4,693 (2.2%) vs 94/4,654 (2.0%); adjusted OR 0.77 (0.55 to 1.08), p = 0.13 |
| Reading: Primary endpoint, not met. Identical result in intention-to-treat and per-protocol analyses. With further adjustment for sex, age, time of day, and day of week, the adjusted OR is 0.72 (0.50 to 1.03; p = 0.07). Note the direction: because the measured event is a failure, an odds ratio below 1 favors the intervention, not against it. The mean risk difference is -0.005 (95% credible interval -0.013 to 0.001), that is, five fewer failures per 1,000 patients treated; the authors bound the plausible effect between 13 fewer and 1 additional failure per 1,000 patients. | |
| Consistency across the 16 sites | Pooled OR 0.76 (95% CrI 0.50 to 1.12); τ = 0.22 (0.01 to 0.67) |
| Reading: Bayesian multilevel meta-analysis on individual-level data. Point estimates point the same way, favoring the intervention, in 14 of 16 facilities, with wide individual intervals. Between-site heterogeneity is low. This is consistency of direction, not a demonstration of efficacy. | |
| Subgroup analyses | No effect modification detected |
| Reading: Prespecified exploratory subgroups: age class (pediatric n = 3,553, adult n = 5,794), day type (weekend n = 561, weekday n = 8,786), sentinel condition (yes n = 1,098, no n = 8,249), and time of consultation (night n = 356, day n = 8,991). The authors write that there is no signal of effect modification across these subgroups. Interaction tests are not corrected for multiplicity and precision is low in several strata. | |
| Clinical documentation quality | Adjusted ORs 1.74; 1.68; 1.71; all p < 0.001 |
| Reading: Across 2,000 consultations reviewed by the panel, assisted clinicians were more likely to record an appropriate diagnosis (adjusted OR 1.74; 1.28 to 2.36), a complete clinical note (1.68; 1.24 to 2.27), and an appropriate treatment plan (1.71; 1.25 to 2.34). This is the trial’s positive result, on a process measure rather than a patient outcome. The authors describe it as strong evidence that the intervention improved documentation and care quality. | |
| Prescribing and sentinel conditions | No difference, except diabetes risk flagging |
| Reading: Correct antibiotic therapy: adjusted OR 0.86 (0.48 to 1.55; p = 0.60). Inappropriate antimalarial prescribing: 0.76 (0.17 to 3.43; p = 0.70). New hypertension diagnosis: 0.85 (0.67 to 1.08; p = 0.20), treatment initiation 0.68 (0.37 to 1.23; p = 0.20). Pediatric acute malnutrition diagnosis: 0.91 (0.50 to 1.64; p = 0.70), referral to a nutritionist 1.14 (0.61 to 2.13; p = 0.70). The only significant difference: fewer patients flagged at risk for type 2 diabetes in the assisted arm, 0.88 (0.78 to 0.98; p = 0.023), which the authors interpret, without demonstrating it, as reclassification of patients whose diabetes was already established. | |
| Satisfaction and consultation duration | Identical satisfaction; median duration 11 minutes in both arms |
| Reading: Among 826 patients who completed the survey, 411 control and 415 assisted, satisfaction scores were identical at the group level (median 4.0; interquartile range 4.0 to 5.0) and the probability of reporting high satisfaction was comparable (adjusted OR 1.02; 0.70 to 1.49; p > 0.9). 784 respondents (95%) judged the duration “just right,” 20 (2.4%) too short, 22 (2.7%) too long. Median consultation duration, a post hoc analysis, was 11 minutes in both arms (interquartile range 7 to 17 minutes in controls, 8 to 17 in the assisted arm; 95% CI of the median difference -1.0 to 0.00; p = 0.031). | |
| Safety | 33 serious events; no causal link retained |
| Reading: 33 serious adverse events across the 16 sites: 27 hospitalizations and 6 deaths. Independent review confirmed appropriate management and found no causal link with the intervention. The post hoc composite of death or hospitalization occurred in 17 controls (0.4%) and 14 assisted patients (0.3%), with no significant difference (adjusted OR 0.77; 0.30 to 1.94; p = 0.60). The authors explicitly write that the trial was not powered to detect rare serious events and that the absence of an observed difference should not be read as proof of safety equivalence. | |
| Quality and adherence to red alerts | 91.8% of outputs judged safe; majority partial adherence |
| Reading: Among 1,000 model outputs flagged red and reviewed by the panel, 494 (49.4%) were judged certainly safe and appropriate and 424 (42.4%) probably safe and appropriate, 40 (4.0%) neutral, 31 (3.1%) probably unsafe or inappropriate, and 11 (1.1%) unsafe and inappropriate. Clinicians fully followed the model’s advice in 195 consultations (19.5%), partially in 573 (57.3%), and not at all in 232 (23.2%). The panel judged the clinician’s decision clinically justified in 284 cases (28.4%) and not justified in 716 (71.6%). This last figure concerns only red-alert consultations and does not generalize to the whole trial. | |
| Costs | Antibiotics: adjusted mean difference -$0.15 (-0.25 to -0.04) |
| Reading: Post hoc analysis. Antibiotics were the largest cost category, US$3.85 per patient in the control arm versus $3.71 in the assisted arm (exchange rate used: $1 to 130 Kenyan shillings); the adjusted mean difference is -$0.147 (95% CI -0.252 to -0.042). This is the only one of twelve drug categories whose confidence interval excludes zero. Mean model cost per patient was $0.04 (95% CI 0.04 to 0.04). | |
Critical appraisal
| Randomization process | Low |
| Variable block randomization at the clinician level, 103 clusters, baseline characteristics comparable across arms for both patients and clinicians. Random-effects models account for the double nesting of clinicians and facilities. The distribution of consultations across facilities is, however, markedly unbalanced between arms, an expected consequence of randomizing practitioners rather than sites, and addressed through adjustment. | |
| Protocol deviations and contamination | Concerns |
| 921 documented deviations and violations: 876 deviations (8.3% of enrolled consultations), mainly non-consented clinicians (570; 5.4%) or multiple clinicians seeing the same consultation (306; 2.9%), 13 consultations that could not be linked to a record (0.1%), and 45 violations (0.4%) from a temporary configuration error that granted control-arm clinicians access to the tool. The authors note that any residual contamination would bias the result toward no difference. Additionally, actual intensity of use was not mandated: the trial measures effectiveness under routine conditions, not performance under enforced use. The authors state this explicitly. | |
| Missing data | Low |
| 11 withdrawals of consent and 90 lost to follow-up out of 9,702 eligible consultations, about 1%. Complete-case analysis, justified by this low proportion. There is an asymmetry, 61 lost to follow-up in the assisted arm versus 29 in the control arm, that is not discussed in the publication. | |
| Outcome measurement | Concerns |
| Blinded adjudication, with double review and an arbiter, is solid. The difficulty lies in the composite itself, which aggregates events of very unequal severity, from a repeat visit for persistent symptoms to death, and can therefore dilute a benefit on any single component. The authors themselves raise the open question of the most appropriate primary endpoint for evaluating a generalist technology across a field as broad as primary care. The 14-day window is short, and the authors list it among their own limitations. | |
| Independence | Concerns |
| Two authors hold stock options in the organization being evaluated, and the model’s provider gave in-kind support to the tool’s development. If anything, these ties would tend to favor a positive result, which makes the absence of a difference on the primary endpoint more credible, not less. Two factors temper further concern: the funder and the model provider are declared to have had no role in the trial’s conduct or publication, and adjudication of clinical endpoints was entrusted to a blinded external panel. | |
| Selective reporting | Low |
| Protocol and statistical analysis plan published ahead of results, primary and secondary endpoints clearly identified, post hoc analyses explicitly labeled as such, non-confirmatory, and not corrected for multiplicity. The results favoring the intervention, documentation and costs, are presented with this caveat by the authors themselves. | |
Level of evidence
Oxford level of evidence 1b. Confidence is high on one point: in this setting, with this tool, at this intensity of use, and within this 14-day window, no reduction in treatment failure was demonstrated, and the trial lacked the precision to rule on modest differences. Confidence is also high on the improvement in clinical documentation quality, which is clear and consistent across all three domains assessed. It is low on everything else: what a more sustained use, a longer follow-up, a specialized rather than generalist tool, or a newer model version would have shown remains unknown, and the authors themselves present their result as a snapshot in time rather than a fixed estimate of current capability. Finally, a trial conducted in a private network of Kenyan clinics says nothing about what would happen in a French psychiatric practice, and it was never meant to.
The colleague test
What an experienced colleague would say if shown this trial in two minutes.
“Finally a randomized trial instead of a vignette demonstration. On the primary endpoint, treatment failure at fourteen days, it finds nothing, and it says so without dressing it up: the interval is wide, the trial was powered for a much bigger effect. What moves is note quality, and that’s a process measure. The interesting part is that the people who ran it had every incentive for a positive result. That said, it’s Kenya, it’s primary care, and fourteen days is short. I’m not changing anything Monday morning, and I’ll note that the case for patient outcomes still has to be made.”
Translation for practice: nothing to change in routine care, and a solid argument against claims made for these tools without a trial behind them.
What you can do with this
- When a decision support tool is pitched to you, the question to ask is about the level of evidence: a comparison on clinical vignettes and a randomized trial in real-world conditions do not answer the same question.
- Always distinguish what is actually measured. Here, clinical documentation improves clearly, while treatment failure at 14 days does not move. A process measure that improves is not a patient benefit, it is a hypothesis of benefit.
- Check the direction of the endpoint. When the measured event is a failure, an odds ratio below 1 favors the intervention. The same 0.77 reads in one direction or the other depending on what is being counted.
- A null result from a team that had an interest in a positive result carries more weight than the same null result from skeptics. The direction of the conflicts of interest should be analyzed, not merely tallied.
- What this trial does not say: it does not say these tools are useless, nor that they are dangerous. The authors write that assistance was safe and did not reduce treatment failure at 14 days, and that a benefit, if it exists, is likely modest. They also write that the trial was not powered to detect rare serious adverse events.
- For a team considering piloting this kind of tool, the critical point raised here is measurement: decide, before deployment, what will count as an improvement, and on which indicators.
- Worth watching: whether equivalent trials exist in mental health was not searched for as part of this verification, which was limited to the source publication and its supplementary material. This analysis can therefore neither confirm nor rule out their existence.
Does this trial show that artificial intelligence makes care worse?
No, and it does not show that it makes care better on this endpoint either. The adjusted odds ratio for treatment failure at 14 days is 0.77, with a 95% confidence interval of 0.55 to 1.08 and p = 0.13. Because the counted event is a failure, the point estimate favors the intervention, but the interval crosses the null value. The authors conclude that assistance was safe, that it did not reduce treatment failure at 14 days, and that a benefit, if it exists, is likely modest.
Is this the first trial of its kind?
The publication does not claim that, and neither does this analysis. The authors write that rigorous evidence in real-world, low-resource settings “remains limited,” which is not the same as nonexistent. They place their work in the continuation of recent randomized evaluations in other clinical contexts, including a randomized trial of a language-model-based chatbot for primary-to-specialist care transitions and a randomized trial of a community-codesigned chatbot in primary care. A quality-improvement study covering 40,000 consultations, using an almost identical tool in the same clinic network, preceded this trial. No claim of priority can be made here.
Why review a Kenyan primary care study in a psychiatry journal?
For the method, not the clinical content. This is one of the few randomized trials to evaluate a generative-AI decision support tool within the real flow of care, with a patient-centered endpoint and independent blinded adjudication. The way it is built, and the way its main result should be read, transfer directly to the day these tools reach psychiatric consultations.
Is a treatment-failure composite a good endpoint?
It has the advantage of capturing several ways a consultation can fail, and the drawback of aggregating events of very unequal severity, from a repeat visit for persistent symptoms to death. The authors themselves raise the open question of the most appropriate primary endpoint for evaluating a generalist technology across a field as broad as primary care, noting that disease-specific endpoints fail to capture the tool’s value, while objective patient outcomes require a scale that is often out of reach.
What would it take to settle this question for psychiatry?
A randomized trial conducted in real psychiatric consultations, with endpoints chosen in advance and meaningful to patients, a measure of actual tool use, and follow-up longer than fourteen days. This verification did not search the literature beyond the source publication, so it cannot say whether such a trial already exists.
Annotated bibliography
Source study. Agweyu A, Mwaniki P, Menon V, Korom R, Isaaka L, Wanyama C, Gill J, Kiptinness S, Adan N, Emmanuel-Fabula M, Riley RD, Archer L, Denniston AK, Liu X, Mateen BA. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine. 2026; volume 32, pages 3032 to 3039. DOI 10.1038/s41591-026-04503-6. Received December 16, 2025, accepted June 2, 2026, published online June 26, 2026. Open access under a Creative Commons Attribution 4.0 license. Trial registration: Pan-African Clinical Trials Registry, number 202502499779176. Funding: Gates Foundation, grant INV-068056; sponsor PATH, Seattle, United States. PMID 42362867. Two supplementary materials were consulted: supplementary tables 1 to 4 and the CONSORT checklist for cluster trials, both fully usable, and the Nature Portfolio reporting summary, obtained by optical character recognition and partially degraded. In that second document, the body text remains legible and was used, but the checklist boxes are illegible and were not used, the model is misrendered as “GPT-40” instead of GPT-4o, and the clinical trial data section, including the registration number, is absent from the excerpt supplied.
Context references. The following references are drawn from the source publication’s bibliography; their identifiers are not present in the sources consulted. Their existence, journal, year, volume, and pagination were confirmed against bibliographic registries on September 2, 2026, with the exception of the Zenodo deposit and the arXiv preprint, which fall outside those registries.
- Menon V et al. Large language model-assisted clinicians versus unassisted clinicians in clinical decision making: protocol for a multi-facility pragmatic cluster randomized controlled trial in Nairobi, Kenya. Zenodo, 2025, https://zenodo.org/records/15788148: the trial’s publicly deposited protocol and statistical analysis plan.
- Agweyu A et al. Safety of a large language model-based clinical decision support system in African primary healthcare. Nature Health. 2026; 1: 607 to 618: a safety evaluation of the same tool by the same team, preceding this trial.
- Korom R et al. AI-based clinical decision support for primary care: a real-world study. Preprint, arXiv 2507.16947, 2025: a quality-improvement study covering 40,000 consultations with an almost identical tool in the same network.
- Tao X et al. An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial. Nature Medicine. 2026; 32: 934 to 942, and Li S et al. A community-codesigned LLM-powered chatbot for primary care: a randomized controlled trial. Nature Health. 2026; 1: 238 to 250: the two recent randomized evaluations the authors describe extending.
- Liu X et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nature Medicine. 2020; 26: 1364 to 1374, and Campbell MK, Piaggio G, Elbourne DR, Altman DG, CONSORT Group. Consort 2010 statement: extension to cluster randomised trials. British Medical Journal. 2012; 345: e5661: the two reporting frameworks followed by the authors.
Regulatory status. Whether such a decision support system would qualify under the EU Medical Device Regulation or the EU Artificial Intelligence Act could not be verified: these texts were not among the sources consulted for this analysis, so no claim is made on this point, which remains to be established by a dedicated review. What the publication does state is that Kenya’s medical device regulator, the Pharmacy and Poisons Board, deemed the product outside its remit, so no local equivalent of an investigational device exemption was filed. The study also states compliance with Kenya’s Data Protection Act of 2019, with personal identifiers removed before any processing by the model.
Editorial collections
Tags
Verified on September 2, 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on September 18, 2026, against the figures of the French version and against the source. How we verify what we publish
This analysis is intended for healthcare professionals. It does not constitute a prescribing recommendation and does not replace individual clinical judgment.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigor.
