Published on 15 September 2026
Adult ADHD: do the treatments hold up when the patient does the rating?
In brief
A component network meta-analysis brings together 113 randomised controlled trials and 14,887 adults with ADHD, to compare 50 treatments dismantled into therapeutic components, pharmacological, psychological and neurostimulatory. Only two components show concordant evidence of benefit across both sources of rating, the patient’s self-report and the clinician’s assessment: stimulants and atomoxetine. For atomoxetine that benefit is paid for by an excess of treatment discontinuation, with an all-cause discontinuation odds ratio of 1.43 (95% CI 1.14 to 1.80) against placebo; extended-release guanfacine does the same, with an odds ratio of 3.70 (95% CI 1.22 to 11.19). The psychological interventions, cognitive behavioural therapy, cognitive remediation, mindfulness and psychoeducation, show a favourable signal when the clinician does the rating, a signal that does not reappear in self-report. Transcranial direct current stimulation follows exactly the same profile. Confidence in the estimates, appraised with the CINeMA framework, ranges from very low to moderate, including for the two main pharmacological results. What is demonstrated here is a hierarchy of evidence, not a hierarchy of true efficacy: the discordance between raters remains a result to interpret, not a verdict on the psychological therapies.
The context
ADHD in adults is rarely treated with a single modality. In the consulting room the question asked is almost never whether methylphenidate works, but whether to add cognitive behavioural therapy, what atomoxetine is worth when stimulants do not suit, whether transcranial stimulation has a place. Those questions call for comparisons between interventions that have almost never been set against each other in the same trial.
Network meta-analysis answers that problem in part, by combining direct and indirect comparisons. Its component variant goes further: it tries to separate, inside combined packages of care, the contribution of each constituent element. That is the main interest of this work, and also its main methodological demand, because dismantling assumes that the components add up, an assumption that can be stated and discussed but not directly verified.
To this is added an old question that is rarely tackled head-on: who judges the improvement. In adult ADHD, self-report and clinician rating do not measure quite the same thing, and blinding is structurally more fragile for a psychological therapy than for a tablet.
What the model measures, and what it does not
A decomposition of effects, not an explanation of mechanism
| Element | Status in the study |
|---|---|
| Effect specific to each component | Estimated by the component model |
| FindingThe estimate rests on an assumption that components are additive. That assumption is a modelling choice, not an observation. If the components interact, the effect attributed to each one is displaced. | |
| Symptoms reported by the patient | Measured, analysed separately |
| FindingThis is the most original contribution of the work: the two sources of rating are not merged into a single outcome, they are handled as two distinct questions. | |
| Symptoms rated by the clinician | Measured, analysed separately |
| FindingComparing the two sets of results is what makes the discordance on the psychological therapies visible. It does not say which of the two sources is wrong. | |
| Tolerability | Approached through discontinuation |
| FindingTreatment discontinuation is a surrogate for tolerability. It aggregates adverse effects, perceived lack of efficacy and practical constraints, and says nothing about the nature of the event. | |
| Quality of life, emotional dysregulation, executive function | Analysed, but as secondary outcomes |
| FindingThese dimensions were indeed analysed, on far smaller networks. On quality of life at 12 weeks, six trials and 1,472 participants only, and no difference demonstrated between the active components and placebo. This is therefore a failure to demonstrate, on an underpowered network, which licenses neither a conclusion of functional benefit nor a conclusion that none exists. | |
The study at a glance
| Population, intervention, comparator, outcome | |
|---|---|
| Population | |
| DetailAdults aged 18 years and over with ADHD diagnosed under DSM-III or a later edition, ICD-10 or ICD-11. 14,887 participants, 113 randomised controlled trials included, retained from 32,416 records identified. Median age 35.6 years. An exclusively adult population, which sets this work apart from syntheses covering children and adults together. | |
| Interventions | |
| Detail50 treatments identified, dismantled into therapeutic components; fourteen active components appear in the main matrix. Medications: stimulants (amphetamines including lisdexamfetamine, methylphenidate), atomoxetine, bupropion, clonidine, extended-release guanfacine, modafinil, viloxazine. Psychological interventions: cognitive behavioural therapy, dialectical behaviour therapy, mindfulness, psychoeducation, cognitive remediation and cognitive training, relaxation. Neurostimulation and neurofeedback: transcranial magnetic stimulation, transcranial direct current stimulation, neurofeedback. Combinations are dismantled into components. | |
| Comparator | |
| DetailPlacebo or another active intervention, through direct and indirect comparisons within the network. | |
| Outcomes | |
| DetailThree primary outcomes: severity of ADHD core symptoms at the timepoint closest to 12 weeks, analysed separately on self-reported and on clinician-reported scales, and acceptability approached through all-cause discontinuation. The numbers contributing differ by outcome: 56 trials and 7,375 participants for the self-reported outcome, 54 trials and 9,742 participants for the clinician-reported outcome, 77 trials and 13,008 participants for acceptability. The legend of figure 1 states 8,083 participants for the self-reported outcome where the body of the text gives 7,375: this internal discrepancy is not explained by the authors. Secondary outcomes: symptoms at 26 and 52 weeks, tolerability, emotional dysregulation, executive function, quality of life. | |
| Design | |
| DetailSystematic review and component network meta-analysis. Random-effects model in a frequentist framework, under R version 4.3.1 with the netmeta package, ranking by p-score. Bayesian models serve only the network meta-regressions on covariates. CEBM level of evidence 1a. Protocol registered with PROSPERO, number CRD42021265576, and published in BMJ Open in 2022. Funded by the National Institute for Health and Care Research, grant NIHR203035, with no industry involvement. |
Quality control
| Domain | Judgement |
|---|---|
| Protocol and registration | Pre-registered |
| FindingPROSPERO registration under number CRD42021265576, protocol published in BMJ Open in 2022, therefore ahead of the analysis. The authors report the departures from the protocol and their reasons in the appendix, which is good practice. One methodological departure is worth noting between the protocol, which announced a Bayesian framework and ranking by SUCRA, and the final report, which adopts a frequentist framework and ranking by p-score. | |
| Search and selection | PRISMA-NMA compliant |
| FindingDuplicate and independent extraction. That is the expected standard for a synthesis of this size. | |
| Risk of bias in the included trials | RoB 2 applied, heterogeneous results |
| FindingThe RoB 2 tool was applied systematically, and separately for each of the primary outcomes. The authors state that they downgraded interventions delivered without blinding, which structurally penalises the psychological arms. Restricting the analysis to trials at low risk of bias alone leaves only 11 trials and 2,281 participants for the self-reported outcome, against 34 trials and 9,480 participants for acceptability. | |
| Confidence in the network estimates | CINeMA, very low to moderate |
| FindingThe framework used is CINeMA, the network version of GRADE logic, and not GRADE in its classical form. It does not apply to the component model, which the authors declare cannot be appraised with CINeMA: confidence is therefore estimated on the comparisons within each sub-network. For the two main pharmacological results, confidence is explicitly placed between very low and moderate. The only high rating reported concerns the acceptability of guanfacine. | |
| Granularity of the network nodes | Stimulants aggregated |
| FindingGrouping the stimulants into a single node, planned in the protocol, limits any differential reading between methylphenidate and amphetamines, and between formulations. The authors ran a sensitivity analysis separating the two families, without demonstrating any notable difference in efficacy. It is a defensible choice for the power of the network, a costly one for the individual decision. | |
| Funding and independence | Public, non-industry |
| FindingPublic funding by the National Institute for Health and Care Research, grant NIHR203035, the funders declaring that they had no role in the conduct of the work nor in the decision to publish. The interests declared outside this work are, by contrast, numerous and involve companies active in the field, notably Angelini Pharma, Takeda, Medice and Janssen. The risk of direct funding bias is low; allegiance bias in method remains possible and cannot be read off a declaration. | |
| Involvement of people with lived experience | At every stage |
| FindingThe authors state that people with lived experience of the condition were involved at every phase, from design and the funding application through to dissemination, and that a patient organisation representative is a co-author. This strengthens the relevance of the outcomes chosen, the self-reported one in particular. | |
The findings
| Intervention | What the data show |
|---|---|
| Stimulants (methylphenidate, amphetamines) | Benefit on both raters, more discontinuation for adverse effects |
| FindingStandardised mean difference −0.39 (95% CI −0.52 to −0.26) on the self-reported outcome and −0.61 (−0.71 to −0.51) on the clinician-reported outcome, at 12 weeks. This is the most solid result of the work. Acceptability, measured by all-cause discontinuation, does not differ from placebo, but tolerability, measured by discontinuation for adverse effects, is worse: odds ratio 2.15 (95% CI 1.61 to 2.88). Grouping into a single node prevents any ranking of the molecules in the primary analysis; a sensitivity analysis separating amphetamines from methylphenidate finds no notable difference in efficacy. | |
| Atomoxetine | Benefit on both raters, excess discontinuation |
| FindingStandardised mean difference −0.38 (95% CI −0.56 to −0.21) on the self-reported outcome and −0.51 (−0.64 to −0.37) on the clinician-reported outcome. All-cause discontinuation odds ratio 1.43 (95% CI 1.14 to 1.80), CINeMA confidence moderate. The interval excludes 1, so the excess of discontinuation is consistent with the data. This indicator does not say why the patient stops: adverse effect, time to onset, or perceived lack of efficacy. Discontinuation for adverse effects is also more frequent, odds ratio 2.43 (95% CI 1.89 to 3.12). | |
| Extended-release guanfacine | Acceptability clearly lower than placebo |
| FindingAll-cause discontinuation odds ratio 3.70 (95% CI 1.22 to 11.19). This is the only result in the whole work carrying high CINeMA confidence, which deserves to be flagged, but the interval is very wide and the estimate rests on few data. No symptomatic benefit is demonstrated for this molecule. | |
| Psychological therapies and interventions | Favourable signal on clinician-reported ratings only |
| FindingCognitive behavioural therapy (−0.76; 95% CI −1.26 to −0.26), cognitive remediation (−1.35; −2.42 to −0.27), mindfulness (−0.79; −1.29 to −0.29) and psychoeducation (−0.77; −1.35 to −0.18) do better than placebo on the clinician-reported scales, and not on the self-reported scales. That absence of signal on self-report does not establish an absence of effect: it establishes that the available data do not allow one to be demonstrated on that outcome. Conversely, relaxation does worse than placebo on the self-reported outcome (0.86; 0.07 to 1.65), an isolated and fragile result. | |
| Neurostimulation and neurofeedback | Signal for tDCS on clinician-reported ratings only |
| FindingTranscranial direct current stimulation does better than placebo on the clinician-reported scales (−0.78; 95% CI −1.13 to −0.43) and not on the self-reported scales: it therefore follows the same discordance profile as the psychological therapies. For transcranial magnetic stimulation and for neurofeedback, no difference is demonstrated on either primary outcome. The whole neurostimulation side rests on ten trials and 194 participants only, a sample that allows neither a benefit to be asserted nor one to be ruled out. | |
| Combinations of interventions | Additivity assumed, judged plausible |
| FindingThe component model postulates that effects are additive, it does not demonstrate it. The authors tested the plausibility of that assumption and conclude that it broadly holds, with a single comparison judged concerning, right-sided transcranial magnetic stimulation against sham stimulation on the self-reported outcome. The real added value of combined packages of care is therefore suggested here, not demonstrated. | |
| Horizon of 26 and 52 weeks | Few data, and a reversed discordance |
| FindingSeventeen trials contribute at 26 weeks and five at 52 weeks, against 87 at 12 weeks. At 52 weeks it is cognitive behavioural therapy, neurofeedback and relaxation that do better than placebo on the self-reported scales, and not on the clinician-reported scales: the gap between raters reverses. On samples that small, that reversal reads as uncertainty, not as a result. | |
| Overall confidence in the evidence | Very low to moderate on the primary outcomes |
| FindingThe CINeMA confidence reported for the pharmacological efficacy estimates ranges from very low to moderate. That is information in itself: even the best supported result of the work does not reach high confidence, and the hierarchy observed reflects the uneven quality of the evidence as much as the true efficacy of the interventions. | |
Critical appraisal
| Point of concern | Judgement |
|---|---|
| Source of the symptom rating | Discordance unresolved |
| FindingThree interpretations remain open and the study does not allow a choice between them: expectation on the part of a clinician rating an intervention they believe in, a real effect that the patient does not perceive subjectively, or inadequacy of the self-report scales in adult ADHD. The three carry opposite clinical consequences. | |
| Blinding of the ratings | Structurally unequal |
| FindingA patient knows whether they are receiving a psychological therapy, and blinding of the rating clinician is harder to guarantee there than in a placebo-controlled trial. Comparing interventions whose blinding differs means comparing estimates whose potential bias is not of the same order. | |
| Transitivity and coherence of the network | Plausible, not verifiable throughout |
| FindingIndirect comparisons assume that the trial populations are close enough to one another. On a network that mixes pharmacology, psychological therapy and neurostimulation, that assumption is more demanding than in a homogeneous network. | |
| Additivity of the components | A modelling choice, not a result |
| FindingThe component model is what gives the work its added value, and what carries its most specific uncertainty. The per-component estimates are to be read as quantified hypotheses. | |
| Precision of the estimates | Good on the pharmacological arms |
| FindingThe pooled numbers are substantial on the pharmacological side. They are far smaller for neurostimulation, where lack of power is enough to explain the absence of a conclusion. | |
| Allegiance bias in the psychotherapy teams | Not formally analysed |
| FindingThe sensitivity analyses reported cover industry funding, duration under 12 weeks, low risk of bias and separation of the stimulants. No dedicated analysis of allegiance bias is reported. This is therefore an absence of analysis and not an absence of signal: that bias can be neither accepted nor dismissed on the strength of this work. | |
| Trial duration and transposition | Short horizon |
| FindingThe data cluster in the short term: 87 trials and 13,680 participants contribute to the analyses at 12 weeks, against 17 trials and 3,282 participants at 26 weeks and five trials and 801 participants at 52 weeks. The longer horizons were indeed analysed, but on networks too thin to settle anything. The authors themselves conclude that their results can reliably guide only the choice of a short-term intervention. | |
| Functional impact | Outside the primary outcome |
| FindingReducing symptom severity is an intermediate outcome compared with what patients ask for: holding down a job, organising a day, keeping a relationship alive. Quality of life was indeed analysed as a secondary outcome, with no difference demonstrated at 12 weeks, but on six trials and 1,472 participants only. The hierarchy established here therefore does not transpose automatically to those goals. | |
Level of evidence
The best established point is this: in this body of trials, stimulants and atomoxetine are the only components whose benefit reappears whatever the source of the rating. The authors put it plainly, writing that “stimulants and atomoxetine were the only interventions with evidence of beneficial effects in terms of reducing ADHD core symptoms in the short term, supported by both self-reported and clinician-reported ratings”. It is a synthesis result, resting on a registered protocol, a standardised risk-of-bias tool, a CINeMA appraisal of confidence and public funding. It has to be noted, however, that the CINeMA confidence attached to those estimates is itself placed between very low and moderate, which forbids talking about it as a high-level result.
Confidence is lower on three points. First on the fine hierarchy between molecules, which the aggregation of stimulants into a single node makes impossible to read in the primary analysis; on that specific question, the dose-response network meta-analysis, which reads adults apart from children, goes further. Next on the non-pharmacological interventions, whose estimates rest on trials that are few, small, and penalised on blinding. Last on the interpretation of the discordance between raters, a finding that is reproducible at 12 weeks but whose explanation remains open, and which reverses at 52 weeks on the few trials available.
What is demonstrated: a hierarchy of the quality of the available evidence. What is suggested: the plausibility of additivity of the components in combined packages of care, tested by the authors and judged broadly tenable. What belongs to expert opinion, including here: the reading one makes of the gap between what the patient says and what the clinician records.
The colleague test
What an experienced colleague would say if you put this study to them in two minutes, between two consultations.
“ On the substance it does not change my prescribing: I still start with a stimulant, and here I have the largest body of trials ever assembled pointing that way. What stops me is the other result. When I do the rating, CBT works. When the patient does the rating, it does not come out. I do not know whether I am the one seeing what I hope to see, or whether the patient is not feeling an effect that is real. Until I know, I am going to stop promising anyone that therapy will make them feel better. I can offer it, I can no longer sell it. ”
What this means in practice: the prescribing strategy does not move, the way the expected benefit of a psychological therapy is announced does. Offering cognitive behavioural therapy as an add-on remains defensible. Telling the patient that they will themselves perceive a reduction in their symptoms is not, for want of data pointing that way on that outcome.
What you can do with this.
- Keep stimulants as first-line treatment in adults, with evidence that is now concordant between self-reported and clinician-reported ratings. State the trade-off at the same time: acceptability does not differ from placebo, but discontinuation for adverse effects is more frequent. The prescribing framework for stimulants is set jurisdiction by jurisdiction, so the rules that apply are the ones in force where you practise.
- Before offering atomoxetine, warn explicitly about the risk of early discontinuation: the all-cause discontinuation odds ratio is 1.43 against placebo. That information changes how the follow-up consultation is prepared. One question comes before it, and its answer is not the same everywhere: the adult indication for atomoxetine differs between jurisdictions, and in France the labelling does not cover initiation in an adult, which places such a prescription outside the marketing authorisation. Check the labelling in force where you practise, along with the cardiovascular contraindications and the interactions.
- Reword what you announce about a psychological therapy: possible benefit, documented above all by the clinician’s observation, not recovered in what the patient reports. Saying it that way is better than a promise the data do not support.
- Do not offer rTMS or tDCS as routine care in adult ADHD on the strength of this work, and present that abstention as uncertainty in the evidence, not as established lack of effect.
- Get into the habit, during follow-up, of collecting both ratings, the patient’s and your own. The study shows that they can diverge. In the consulting room, that divergence is clinical information in itself.
- The course of action is set out in the NICE decision tree for depression in adults.
Frequently asked questions
Does this study mean that CBT is useless in adult ADHD?
No. It shows that the favourable signal seen on clinician-reported ratings at 12 weeks is not recovered on self-report. A failure to demonstrate is not proof that there is no effect. The cautious conclusion is that subjective benefit is not demonstrated on that outcome, not that it is nil. It is worth noting, moreover, that at 52 weeks, on the five trials available, cognitive behavioural therapy comes out on self-report and not on clinician rating, which is an invitation not to fix the interpretation.
Should we conclude that stimulants are superior to everything else?
They are the component whose benefit is best supported, which is not the same thing. Part of the gap comes from the uneven quality of the evidence: the pharmacological trials are more numerous, larger, and allow firmer blinding. The primary analysis does not rank the molecules against one another, since the stimulants form a single node of the network; a sensitivity analysis separating amphetamines from methylphenidate demonstrated no notable difference in efficacy. Finally, stimulants lead to more discontinuation for adverse effects than placebo, with an odds ratio of 2.15.
Should a discontinuation odds ratio of 1.43 rule out atomoxetine?
No, it should make you prepare the prescription. The benefit on symptoms is recovered on both sources of rating. The excess of discontinuation is real in the data, but all-cause discontinuation aggregates very different reasons, adverse effects, time to onset, the constraints of daily life. That is an argument for anticipating and monitoring, not for abstaining. One question comes before this one and it is regulatory: the adult indication for atomoxetine is not the same in every jurisdiction, so check the labelling in force where you practise.
A patient asks me about transcranial stimulation, what should I say?
That a signal exists, and that it is fragile. Transcranial direct current stimulation does better than placebo on the scales rated by the clinician, without that being recovered in what the patient reports, and the whole neurostimulation side rests on ten trials and 194 participants only. For transcranial magnetic stimulation, no difference is demonstrated. So this is not a refusal grounded in proof of lack of effect, it is an abstention grounded in insufficient evidence. The distinction matters for the therapeutic relationship.
Is a network meta-analysis worth a head-to-head trial?
It does not replace one. It estimates comparisons that have never been run, from an assumption of similarity between the trial populations. The component model adds an assumption of additivity. Those assumptions are reasonable, they are not directly verifiable, and they justify reading the network rankings as orders of magnitude, not as measurements.
Annotated bibliography
Source study. Ostinelli EG, Schulze M, Zangani C, Farhat LC, Tomlinson A, Del Giovane C, Chamberlain SR, Philipsen A, Young S, Cowen PJ, Bilbow A, Cipriani A, Cortese S. Comparative efficacy and acceptability of pharmacological, psychological, and neurostimulatory interventions for ADHD in adults: a systematic review and component network meta-analysis. The Lancet Psychiatry 2025; 12(1): 32-43. DOI 10.1016/S2215-0366(24)00360-2 · PMID 39701638. PROSPERO registration CRD42021265576. The title carried by the present analysis is an editorial title, distinct from the published title recalled above. Contribution: the largest comparative synthesis of interventions for adult ADHD, with the two sources of rating analysed separately. Limitation: aggregation of the stimulants into a single node in the primary analysis, and CINeMA confidence placed between very low and moderate, including for the main pharmacological estimates.
Context reference 1. Cortese S, Adamo N, Del Giovane C, et al. Comparative efficacy and tolerability of medications for attention-deficit hyperactivity disorder in children, adolescents, and adults: a systematic review and network meta-analysis. The Lancet Psychiatry 2018; 5(9): 727-738. DOI 10.1016/S2215-0366(18)30269-4 · PMID 30097390. Contribution: the earlier pharmacological reference, cited by the authors as a point of comparison for their own results. Limitation: pharmacological scope only, with paediatric and adult populations handled in the same work, which makes direct comparison with the present network imperfect.
Context reference 2. Faraone SV, Banaschewski T, Coghill D, et al. The World Federation of ADHD International Consensus Statement: 208 evidence-based conclusions about the disorder. Neuroscience and Biobehavioral Reviews 2021; 128: 789-818. PMID 33549739. Contribution: a reference frame on the validity of the diagnosis and on the epidemiological data. Limitation: a consensus document, therefore structured expert opinion and not a quantitative synthesis, with numerous declared interests to examine.
Context reference 3. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Medicine 2020; 17(4): e1003082. DOI 10.1371/journal.pmed.1003082 · PMID 32243458. Contribution: sets out the method used here to appraise confidence, and therefore the real reach of the network estimates. Limitation: a methodological tool whose application involves an element of judgement, leaving variability between teams; the authors note in addition that it does not apply to the component model itself.
Context reference 4. National Institute for Health and Care Excellence. Attention deficit hyperactivity disorder: diagnosis and management. Guideline NG87, published in March 2018, last updated in September 2019. Contribution: a practice framework setting out the respective places of drug treatments and non-drug interventions in adults, and cited as such by the authors. Limitation: written for the British care system, and predating the present synthesis, which the authors propose precisely to use in revising it.
