Published on 2 October 2026

Analysis · Psychosis and bipolar disorder · Precision psychiatry

Joint prediction of psychosis and bipolar disorder risk from the clinical record: what is the South London model worth?

▤ Dossier The Lancet Psychiatry · 2026 · 13(1): 14-23 · Arribas et al. DOI: 10.1016/S2215-0366(25)00307-4 PMID 41317739 Scientific 76 Editorial 78

The essentials

The authors built a model for patients in secondary mental health care who receive a first diagnosis of a mental disorder that is neither psychotic nor bipolar. It estimates their risk of being diagnosed with a psychotic disorder or a bipolar disorder within six years. The model uses only data already held in the electronic health record: age, gender, ethnicity, index diagnosis, prescriptions, hospital stays, and signs, symptoms and substance use extracted from clinical notes by natural language processing. It was developed and evaluated in 127,868 patients of the South London and Maudsley NHS Foundation Trust (SLaM) using internal-external validation: trained on four of the trust’s five geographical areas and tested on the fifth, in turn. Across the left-out areas, the mean concordance index is 0.80 (95% CI 0.78 to 0.81) and the mean calibration slope 1.02 (SD 0.14). Decision curve analysis shows a net benefit over default strategies across the whole range of thresholds studied (0 to 50%), with gains concentrated at low thresholds and an overall utility that the authors themselves describe as modest. The main limitation is plain: one trust, one city, no independent external validation. As it stands, the model cannot be used outside the service where it was built, in France for example.

Context

For the authors, identifying people at risk of psychosis or bipolar disorder early is a precondition for effective prevention, and current detection strategies remain inefficient. The two disorders overlap considerably, in their prodromal phase as in their presentation: psychotic symptoms affect 63% of patients with bipolar I disorder over their lifetime (95% CI 58 to 68), and 16% of people at clinical high risk for psychosis also meet criteria for clinical high risk for bipolar disorder.

Existing prediction models were developed separately for each disorder. Only one joint model had been published, and it performed poorly on external validation (area under the ROC curve 0.64, and 0.62 for bipolar disorder). This study sets out from what the electronic record collects anyway, and puts it to work to identify the patients who warrant a closer assessment.

The study at a glance

Item
Population
Patients of all ages (mean age 33.4 years, SD 18.8) who received at the South London and Maudsley NHS Foundation Trust, between 1 January 2008 and 10 August 2021, a first ICD-10 diagnosis of a non-organic, non-psychotic and non-bipolar mental disorder, or a designation of clinical high risk for psychosis; exclusion of index dates falling within the initial window from 1 January to 30 June 2008, of patients receiving clozapine or a long-acting injectable antipsychotic, and of patients with no data after the index date: 127,868 patients in total
Predictors
Routine electronic health record data, recorded at the index date or in the six months before it: 77 candidate predictors, including 5 sociodemographic and clinical (age, gender, age-by-gender interaction, ethnicity, index diagnosis), 4 medication (antipsychotics and mood stabilisers below the minimum effective dose, antidepressants, anxiolytics and hypnotics), 2 hospitalisation, and 66 signs, symptoms and substance use variables extracted from clinical notes by natural language processing. Final model: 28 predictors. Family history not available
Samples
Internal-external validation: model trained in turn on four of the trust’s five areas (the boroughs of Croydon, Lambeth, Lewisham and Southwark, and patients referred from the rest of London) and tested on the fifth; performance pooled by random-effects meta-analysis; final model fitted on all 127,868 patients
Outcome
First ICD-10 diagnosis recorded in the record of a psychotic disorder (F20 to F29 excluding F21 and F23, manic, bipolar or depressive episodes with psychotic features, puerperal psychosis, substance-induced psychotic disorders) or a bipolar disorder (F30 and F31 without psychotic features, cyclothymia F34.0) within six years; discrimination (Harrell’s C-index), Brier score, calibration, clinical utility by decision curve analysis
Design
Retrospective prognostic cohort study using an electronic health record registry; LASSO-regularised Cox proportional hazards model; single imputation of missing covariates by random forests; internal-external validation by geographical area; mean follow-up 622 days (SD 687), median 315 days (interquartile range 75 to 1,005)

Quality check

ItemJudgement
Sample sizeVery large
Finding127,868 patients and 3,150 events. According to the authors’ sample size calculation, the data allowed up to 500 parameters, against about 90 candidate parameters before LASSO selection.
ValidationInternal-external only
FindingThe model is tested on geographical areas of the same trust, each left out in turn. No validation in another service or another country.
Statistical methodAppropriate
FindingLASSO penalisation against overfitting, internal-external validation, calibration reported alongside discrimination, decision curve analysis. Competing risks between psychosis and bipolar disorder were not modelled: according to the authors, no available tool allowed them to be combined with LASSO penalisation.
TransparencyStated
FindingReported in line with the RECORD and TRIPOD+AI guidelines, analysis code publicly available. The record data remain behind the NHS firewall, with no permission to share them, which prevents direct independent reproduction.
Funding and competing interestsPublic funders
FindingFunding from the UK Medical Research Council and the NIHR Biomedical Research Centres (South London and Maudsley, Oxford Health); the funders had no role in the study. Industry competing interests declared by five of the twelve authors, including employment at F Hoffmann-La Roche for the first author, outside this study.

Results

0.80
Mean concordance index of the model across the areas left out in internal-external validation (95% CI 0.78 to 0.81), in 127,868 patients.
ItemValue
Discrimination, by area0.79 to 0.82
FindingHarrell’s concordance index in each of the five areas tested: 0.79 in Croydon, Lambeth and Lewisham, 0.80 in Southwark, 0.82 for patients referred from the rest of London. Overall performance in the development sample is not reported.
Discrimination, validation0.80 (95% CI 0.78 to 0.81)
FindingMean across the five left-out areas, pooled by random-effects meta-analysis. Performance varies little from one area to another.
CalibrationSlope of 1.02 (SD 0.14)
FindingA slope close to 1 means that predicted risks are, on average, neither overestimated nor underestimated. It nevertheless ranges from 0.90 to 1.23 across areas. Calibration-in-the-large: 0.06 (SD 0.02); six-year Brier score: 0.03 (SD 0.01).
Clinical utilityNet benefit for thresholds of 0 to 50%
FindingAcross this range, using the model does better than assessing everyone or no one. At the 10% threshold, the net benefit is 3%, which the authors equate to three additional cases detected per 100 patients screened. Gains are concentrated at low thresholds, and the authors describe overall clinical utility as modest.
Six-year cumulative incidence8.27% (95% CI 7.84 to 8.70), psychosis or bipolar disorder
FindingCombined incidence of the two disorders; separate cumulative incidences are not reported. The sample includes 3,150 events (2.5%). The outcome remains uncommon: even a good model can produce many false positives.

Critical appraisal

ItemJudgement
External validityNot established
FindingOne trust in South London, with a highly diverse urban population. The stability of performance across the five areas (concordance index 0.79 to 0.82) is a first sign of transportability within the trust; nothing yet shows that the model keeps its performance elsewhere, and the authors do not expect it to outside secondary care.
Diagnostic codingRecord-dependent
FindingThe outcome is a recorded diagnosis, not one confirmed by structured interview.
Prescription as predictorConfounding by indication
FindingThe prescriptions retained in the model (mood stabilisers, anxiolytics and hypnotics, antidepressants), like the symptoms extracted from the notes, partly reflect the clinician’s judgement. The model therefore also learns what doctors already suspected. Antipsychotics, for their part, were dropped by LASSO selection.
Authors’ claimsOptimistic, limitations acknowledged
FindingThe authors describe the performance as excellent and consider the model feasible in practice at the 10% threshold. They acknowledge that independent external validation and prospective trials are still needed, and that the model does not apply outside secondary care. Their claim of priority is a narrow one: the first joint model for the two disorders “with excellent performance”, an earlier joint model having existed.
What is not measuredEffect on patients
FindingThis study does not test whether using the model brings diagnosis forward or improves patient outcomes; the authors consider prospective effectiveness trials necessary.

Level of evidence

Scientific76/100
Editorial78/100

The level of evidence is that of a retrospective cohort prognostic study (Oxford CEBM level 2b). Confidence is good in the performance measured within this trust: a very large sample, internal-external validation on distinct geographical areas, and both calibration and discrimination reported.

It is low as soon as one leaves South London. External validation in another health system is the next step, and nothing replaces a trial showing a concrete benefit for the patients the model flags.

The colleague test

What an experienced colleague would say if you presented this study in two minutes, between two consultations.

« It’s serious work, and it shows that psychosis and bipolar disorder can be flagged together with what is already in the record. But it’s a single London trust. Where I practise, nothing like it exists. What I take from it is to structure the assessment of young people who arrive with vague symptoms more carefully. »

Translation for practice: no tool to install, but a list of signals not to let slip in a young patient seen for the first time.

What you can do with this

  • What you can understand: risk of psychosis or bipolar disorder can partly be read in record data collected at the first diagnosis, and, in this study, a joint model for the two disorders did not perform significantly worse than two separate models.
  • What you can do at a young patient’s first appointment: record systematically substance use, psychotic symptoms (persecutory ideas, hallucinations) and previous psychotropic prescriptions; family history, absent from this model, remains an important risk factor according to the authors.
  • What you can tell a patient or their family: no calculation says who will develop a disorder; a model like this one serves to decide whom to see again more closely, not to make a diagnosis.
  • What you can watch for: an external validation of the model, particularly outside the United Kingdom, and an impact study on time to diagnosis.
  • What you can teach: the difference between discrimination, calibration and clinical utility, all three of which this study reports.

Frequently asked questions

Can this model be used outside the trust where it was built?

Not as it stands. It was built and tested in a single NHS trust, with its own coding and organisation of care. It would need to be validated on local data before any use elsewhere: in France, for example, on French data.

What does a concordance index of 0.80 mean?

Take at random one patient who will receive the diagnosis and one who will not: the model assigns the first a higher risk in about 80% of cases. That is good discrimination for a clinical model, but on its own it says nothing about how useful the model is in the consulting room.

Why predict psychosis and bipolar disorder together?

Because the two disorders share some of their early signs and clinical features. A joint model aims to avoid missing a patient because the clinician was looking for the other disorder; this benefit was not measured in the study.

Does the model use laboratory tests or imaging?

No. It relies only on information already in the clinical record. That is its practical appeal, and it is also what makes it dependent on the quality of that record.

Will a patient classed as high risk necessarily develop a disorder?

No. The six-year cumulative incidence of either disorder is 8.27% in this cohort. The paper does not report the proportion of patients classed as high risk who actually developed a disorder: a predicted risk remains a probability, not an individual prognosis.

Annotated bibliography

Source study. Arribas M, de Micheli A, Krakowski K, Stahl D, Correll CU, Young AH, Andreassen OA, Vieta E, Arango C, McGuire P, Oliver D, Fusar-Poli P. Joint detection of risk for psychotic disorders or bipolar disorders in clinical practice in the UK: development and validation of a clinical prediction model. The Lancet Psychiatry. 2026;13(1):14-23. Published online 26 November 2025. DOI: 10.1016/S2215-0366(25)00307-4. PMID: 41317739. Funding: UK Medical Research Council (MR/N013700/1), National Institute for Health Research (NIHR) Biomedical Research Centres at South London and Maudsley NHS Foundation Trust, and Oxford Health NHS Foundation Trust; the funders had no role in study design, data collection, analysis, interpretation or writing. Declared competing interests: MA has been employed by F Hoffmann-La Roche, outside the study; AHY has served as a consultant and advisor for Flow Neuroscience, Novartis, Roche, Janssen, Takeda, Noema Pharma, Compass, AstraZenaca, Boehringer Ingelheim, Eli Lilly, LivaNova, Lundbeck, Sunovion, Servier, Livanova, Janssen, Allegan, Bionomics, Sumitomo Dainippon Pharma, Sage and Neurocentrx; OAA has been a consultant to Cortechs.ai and Precision Health and has received speaker’s honoraria from Lilly, Janssen, Otsuka and Lundbeck; EV has received grants and served as consultant, advisor or CME speaker for AB-Biotics, AbbVie, Adamed, Alcediag, Angelini, Biogen, Beckley-Psytech, Biohaven, Boehringer Ingelheim, Casen-Recordati, Celon Pharma, Compass, Dainippon Sumitomo Pharma, Esteve, Ethypharm, Ferrer, Gedeon Richter, GH Research, GlaxoSmith Kline, HMNC, Idorsia, Johnson & Johnson, Lundbeck, Luye Pharma, Medincell, Merck, Newron, Novartis, Orion Corporation, Organon, Otsuka, Roche, Rovi, Sage, Sanofi-Aventis, Sunovion, Takeda, Teva and Viatris, outside the submitted work; PF-P has received research funds or personal fees from Lundbeck, Angelini, Menarini, Sunovion, Boehringer Ingelheim, Mindstrong and Proxymm Science, outside the study; the remaining authors (AdM, KK, DS, CUC, CA, PM and DO) declare no competing interests.

Editorial collections

Tags

Verified on 2 October 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on 2 October 2026, against the figures of the French version and against the source. How we verify what we publish

This analysis is intended for healthcare professionals. It does not constitute a prescribing recommendation and does not replace individual clinical judgment.

Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigor.

Report an error in this analysis

Follow Dr Stroescu on LinkedIn, for the review every Saturday