Published on 16 September 2026

Analysis · Depression · Psychopharmacology

Treatment-resistant depression: what is a network meta-analysis ranking of treatments worth?

★ Premium Neuropsychopharmacology · 2025; 50 (6): 913-919 (published online 30 December 2024) · Saelens et al. DOI 10.1038/s41386-024-02044-5 PMID 39739012 Scientific 76 Editorial 90

In brief

This network meta-analysis brings together 69 randomised trials and 10,285 adult participants with treatment-resistant depression, strictly defined as failure of at least two antidepressants. Of twenty-five treatments compared with placebo or sham, only six come out ahead on response rate: electroconvulsive therapy, minocycline, theta-burst stimulation, repetitive transcranial magnetic stimulation, ketamine and aripiprazole. The resulting hierarchy runs against intuition, since it places neuromodulation and ketamine ahead of antipsychotic augmentation. It is also far more fragile than its top position suggests: the odds ratio of 12.86 attributed to electroconvulsive therapy comes with a 95% confidence interval running from 4.07 to 40.63, and that superiority disappears as soon as trials at high risk of bias, unblinded trials, trials with imputed data or trials without a placebo condition are removed. To be read as a pointer to direction, not as a league table.

The context

The question comes up in every second-opinion consultation. A patient has not responded to two adequately conducted antidepressant trials, and the next step has to be chosen. The options are many and rarely compared with one another within a single trial: switching class, augmenting with an antipsychotic or with lithium, referring for ketamine, for magnetic stimulation, for electroconvulsive therapy. Existing trials almost always compare one product with a placebo, never two strategies with each other. Network meta-analysis is the tool that makes it possible to reconstruct these missing comparisons by going through the common comparator.

The methodological point

What the method allows, and what it assumes

A network meta-analysis estimates an effect between two treatments that have never been compared directly, by going through the comparators they share. The price of this extension is an assumption, transitivity: the populations, outcomes and settings of the linked trials must be similar enough for the indirect comparison to make sense. The authors state that they checked it on several covariates and that they duplicated the frequentist analysis with a Bayesian analysis. That is the right level of rigour. It does not remove the need to look at what feeds each node of the network, because a top ranking can rest on three trials and two hundred patients.

The study at a glance

Population
Adults aged 18 and over with treatment-resistant depression, defined as lack of response to at least two antidepressants, following the definition endorsed by the FDA and the EMA. Bipolar disorders and psychotic depression excluded. Mean age 43.73 years (SD 11.29), 5,662 women (55.05%), mean baseline MADRS score 33 (SD 6), mean duration of the current episode 33 months (SD 38.64).
Interventions
Twenty-five treatments: antidepressants, antipsychotic augmentation, mood stabilisers, ketamine and psychedelics, neuromodulation (electroconvulsive therapy, repetitive transcranial magnetic stimulation, theta-burst stimulation, transcranial direct current stimulation, deep brain stimulation).
Comparator
Placebo and sham, pooled into a single node. This pooling is a methodological choice whose consequences are discussed below.
Outcomes
Primary, response rate, defined as a reduction of at least 50% in the depression score, on scales that varied across trials. Secondary, remission, depression score at endpoint and tolerability, the latter defined as the proportion of withdrawals from the trial due to adverse events.
Design and numbers
Systematic review and network meta-analysis of randomised trials, searching PubMed, Embase and the Cochrane CENTRAL register, restricted to English, from inception of the databases to 13 April 2023. 8,234 records screened, 390 full texts examined, 69 trials included, 10,285 participants. Network for the primary outcome, 67 trials and 9,354 participants, two trials not providing enough to calculate a response rate. Mean duration of the included trials, 5.07 weeks.

Quality control

CriterionJudgement
Design and internal replicationSound
FindingFrequentist analysis (Netmeta package) duplicated within a Bayesian framework, strict regulatory definition of resistance, transitivity examined on baseline severity, duration of the episode, comorbidities, concurrent treatments, age and sex.
Quality of the included trialsFragile base
FindingIn the Results section, 42 trials (60.9%) are rated at moderate risk of bias, 9 (13%) at high risk and 18 (26.1%) at low risk, with the Cochrane RoB 2 tool. The abstract gives a different breakdown (12.5% high risk, 59.38% some concerns, 28.13% low risk), an inconsistency the authors do not explain. A well-conducted review on a fragile base remains a review on a fragile base.
Precision at the top of the rankingMajor imprecision
FindingElectroconvulsive therapy rests on 208 patients, minocycline on a single study of 41 patients, whereas aripiprazole counts 1,933 and repetitive magnetic stimulation 1,810. The ranking is built on P-scores (0.96 for electroconvulsive therapy on response rate), which summarise estimates some of which come from a single trial.
Sensitivity analysesFragile results
FindingElectroconvulsive therapy is no longer significant against sham when trials with imputed data, those at high risk of bias, unblinded trials or those without a placebo condition are excluded. Minocycline is no longer significant in non-commercially funded trials only. For aripiprazole this analysis is impossible: the six included trials are all industry-funded, so the lack of significance there reflects an absence of data, not a disappearance of effect.
Pooling of placebo and shamDebatable
FindingThe mean response rate is 22.48% in placebo arms (range 0 to 49.15%) and 14.07% in sham conditions (range 0 to 72.72%). Separating the two nodes gives results close to the main analysis, with one exception: brexpiprazole then reaches significance (OR 1.53, 95% CI 1.03 to 2.27). The pooling assumption therefore weighs on at least one result.
Heterogeneity and publication biasModerate
FindingHeterogeneity I² 47.3%, 95% CI 26.8 to 62%. Significant between-design heterogeneity (Q 18.53, p 0.018) and within-design heterogeneity (Q 76.27, p 0.001); under the full design-by-treatment interaction model, between-design inconsistency is no longer significant (Q 6.79, p 0.56). No indication of publication bias was noted by the authors, on comparison-adjusted funnel plots and an Egger test.
TransparencyGood, with reservations
FindingInstitutional funding (Intramural Research Program of the National Institute of Mental Health, ZIAMH002857, open access publication funded by the Medical University of Vienna), competing interests declared, compliance with PRISMA and its extension for networks, PROSPERO registration. Two reservations: data not included are available only on request from the corresponding author, and the registration number differs between the abstract (CRD42023420584) and the methods section (CRD42022324095).
Authors’ competing interestsTo be taken into account
FindingCarlos A. Zarate Jr is a co-inventor on patents covering the use of ketamine and its metabolites in depression, with rights assigned to the US government and a share of any royalties. Christoph Kraus and Rupert Lanzenberger declare honoraria, research funding or support from several companies, including the one that markets esketamine. The other authors declare no competing interests. This does not invalidate the ketamine result, but it belongs in the file when reading it.

The findings

6 / 25
Six treatments out of twenty-five come out ahead of the comparator on response rate. Nineteen are not distinguishable from it, which is not the same thing as being ineffective.
OutcomeReported value
Electroconvulsive therapyOR 12.86, 95% CI 4.07 to 40.63
PEB readingTop of the ranking, but the interval spans a ratio of one to ten. The network contains no trial of electroconvulsive therapy against placebo or sham: the estimate comes from three head-to-head comparisons against repetitive magnetic stimulation or transcranial direct current stimulation. The authors point out that “there has been no sham-controlled study of ECT for either TRD or non-TRD since 1985”.
MinocyclineOR 6.5, 95% CI 1.27 to 33.29
PEB readingA node fed by a single study of 41 patients, which the authors report stands out for an unusual baseline severity (mean HAM-D score 34.5). This result does not carry over to clinical practice.
Theta-burst stimulationOR 4.8, 95% CI 2.21 to 10.39
PEB readingA clear effect, a wide interval but lying entirely above 1, in 683 participants.
Repetitive magnetic stimulationOR 4.01, 95% CI 2.36 to 6.81
PEB readingThe best-supported result among the top-ranked treatments, with the narrowest interval and 1,810 participants.
KetamineOR 3.4, 95% CI 2.14 to 5.41
PEB readingIn the subgroup analyses, only the racemic form separates significantly from placebo, but the direct comparison between racemic ketamine and esketamine, routes of administration pooled, shows no significant difference. The TRANSFORM trials, which the authors describe as “pivotal for the approval of (S)-ketamine for TRD”, were excluded from the main analysis on transitivity grounds, an oral antidepressant being newly started in both arms.
AripiprazoleOR 1.9, 95% CI 1.25 to 2.91
PEB readingThe only antipsychotic to stand out, in 1,933 participants. Main reservation: the six included aripiprazole trials are all industry-funded, which rules out testing this result on independent trials. This is not a demonstration of an effect driven by funding, it is the absence of any data that would allow one to be excluded.
Effects by broad classNeuromodulation 3.35, NMDA targets 2.94, antipsychotics 1.36
PEB readingNeuromodulation OR 3.35 (95% CI 2.09 to 5.35), agents targeting the NMDA receptor OR 2.94 (2.03 to 4.27), antipsychotics OR 1.36 (1.04 to 1.78). The other three classes do not separate from placebo: serotonergic psychedelics 2.26 (0.91 to 5.61), other pharmacological mechanisms 1.44 (0.86 to 2.39), mood stabilisers 1.09 (0.63 to 1.88).
Not significantPsilocybin 1.8, quetiapine XR 0.99, lithium 0.71
PEB readingPsilocybin OR 1.8 (95% CI 0.62 to 5.25), extended-release quetiapine 0.99 (0.48 to 2.08), lithium 0.71 (0.28 to 1.83). No difference demonstrated in this network, which does not demonstrate an absence of effect. For psilocybin, the number of trials available at the date of the search remains small.
Secondary outcomesRemission, endpoint score, tolerability
PEB readingFour treatments hold up on every outcome: electroconvulsive therapy, theta-burst stimulation, repetitive magnetic stimulation and ketamine. Remission is significantly different from placebo for these four treatments, plus aripiprazole and the olanzapine and fluoxetine combination. On tolerability, six treatments significantly increase withdrawals from the trial due to adverse events (odds ratios from 2.23 to 5.63): quetiapine, olanzapine, fluoxetine, the olanzapine and fluoxetine combination, aripiprazole and brexpiprazole. Antipsychotics are the least well tolerated class in the analysis.

Critical appraisal

DomainJudgement
Validity of the question askedHigh
FindingRegulatory definition of resistance, exclusion of bipolar and psychotic forms: the population is that of the real clinical problem.
Internal validity of the networkModerate
FindingThe review is better conducted than the trials it pools. The quality of a synthesis does not make up for the quality of its raw material.
Fit between claim and evidenceNeeds qualifying
FindingThe authors’ conclusions remain measured, they write: “These findings may help guide evidence-based treatment choices for TRD”. But the abstract highlights a range of odds ratios reaching up to 12.86, without pairing it with the sensitivity analyses that bring that figure down. The hurried reader will remember the 12.86.
External validityPartial
FindingThe outcome is short-term response, in trials with a mean duration of 5.07 weeks, rarely beyond 6 to 8 weeks. Maintenance of response, long-term tolerability and acceptability to the patient are not what this network ranks, and the authors explicitly acknowledge it.
Claim of priorityCircumscribed
FindingThe authors do not claim the first network meta-analysis of treatment-resistant depression: they cite the earlier ones. They claim to be “the first to incorporate neuromodulatory treatments alongside both established antidepressants and novel, rapid-acting antidepressants”, and the first to include serotonergic psychedelics in such a network. The restriction is their own and is reproduced as it stands.
Documentation of resistanceSelf-reported
FindingParticipants had on average failed 4.33 antidepressants (SD 1.96), but only 17 trials (24.64%) report the exact number of previous treatment trials. The regulatory definition is applied to data that are often reported and not verified.

Level of evidence

Scientific76
Editorial90

Oxford level of evidence 1a by design: systematic review with network meta-analysis of randomised trials. The level reflects the design, not the quality of the material, and nearly three quarters of the pooled trials are rated at moderate or high risk of bias. Confidence is moderate that neuromodulation and agents targeting the NMDA receptor sit above antipsychotic augmentation in this population, a result carried by the class-level analysis. It is low on the exact order of the treatments among themselves, and very low on the top two positions taken in isolation. This dissociation between confidence in the message and distrust of the ranking is the methodological lesson of the article.

The colleague test

What an experienced colleague would say if you put this study to them in two minutes, between two consultations.

“ In genuine resistance, two documented failures, I think neuromodulation and ketamine before stacking up antipsychotics, where only aripiprazole stays afloat. But I do not take the first place of electroconvulsive therapy at face value: it rests on 208 patients and three head-to-head comparisons, with not a single trial against sham in the network. ”

What this means in practice: what remains from this work is an order of magnitude between classes of treatment, not a league table. The decision remains an individual one, and it takes in access, tolerability and the patient’s preference, none of which appear in any of these odds ratios.

What you can do with this

  • First check that the resistance is real: two antidepressants, at an adequate dose, for an adequate duration, with documented adherence. The network applies only to that population.
  • Consider referral for neuromodulation or ketamine earlier in the care pathway, rather than after a third or fourth layer of medication.
  • If antipsychotic augmentation is chosen, bear in mind that aripiprazole is the only one to stand out in this network, that its six included trials are all industry-funded, with no independent check available, and that antipsychotics are the least well tolerated class in the analysis.
  • Distinguish racemic ketamine from esketamine: in this network only the racemic form separates from placebo, whereas the direct comparison between the two forms shows no significant difference. Their access conditions differ, and they vary between health systems. In France, for example, intranasal esketamine holds a marketing authorisation in treatment-resistant depression, in combination with a serotonergic antidepressant, with administration under supervision and a contraindication in cerebrovascular disease, aneurysm or arteriovenous malformation, or a history of intracerebral haemorrhage. There, racemic ketamine remains outside any authorisation for this indication.
  • Bear in mind that, in France for example, two of the six winning treatments have no indication in depression: aripiprazole is not authorised there as augmentation of an antidepressant, and minocycline has no psychiatric indication there. In that setting, these are research data, not options for immediate prescription.
  • Be ready to answer a patient who has read that electroconvulsive therapy is supposedly twelve times more effective: the figure exists, the interval that comes with it runs from 4.07 to 40.63, and it does not survive the removal of the weakest trials.

Frequently asked questions

Should the order of treatments in treatment-resistant depression change?

The work suggests not holding back neuromodulation and ketamine as a last resort. It does not provide an algorithm, and it does not compare sequential strategies with one another.

Why does lithium augmentation not appear effective in treatment-resistant depression?

Its estimate in this network is not distinguishable from the comparator (OR 0.71, 95% CI 0.28 to 1.83). That does not demonstrate that it is ineffective. The authors themselves offer an explanation: earlier meta-analyses favourable to lithium defined resistance as failure of a single antidepressant and included bipolar depression, two situations excluded here.

Does this treatment ranking apply to bipolar depression?

No. Patients with bipolar disorder were explicitly excluded, as were those with psychotic depression.

Is psilocybin effective for treatment-resistant depression?

Psilocybin does not reach significance in this network (OR 1.8, 95% CI 0.62 to 5.25), any more than ayahuasca does (OR 3.94, 95% CI 0.73 to 21.20). The literature search ends on 13 April 2023 and the authors consider further studies necessary to settle the question. No firm conclusion, in either direction, can be drawn from it.

Is an odds ratio of 12.86 for ECT credible?

It matches what the authors report, but it is very imprecise and it does not withstand the sensitivity analyses. It is a textbook case: the size of an effect and the strength of the evidence are two different things.

Annotated bibliography

Source study. Saelens J, Gramser A, Watzal V, Zarate CA Jr, Lanzenberger R, Kraus C. Relative effectiveness of antidepressant treatments in treatment-resistant depression: a systematic review and network meta-analysis of randomized controlled trials. Neuropsychopharmacology. 2025;50(6):913-919. Published online 30 December 2024. DOI 10.1038/s41386-024-02044-5. PMID 39739012. Open access version: PMC12032262. Network meta-analysis of 69 randomised trials, institutional funding from the National Institute of Mental Health, extensive supplementary material hosted by the publisher.

Direct counterpoint. Anand A, Mathew SJ, Sanacora G, Murrough JW, Goes FS, Altinay M, et al. Ketamine versus ECT for nonpsychotic treatment-resistant major depression. N Engl J Med. 2023;388(25):2315-2325. DOI 10.1056/NEJMoa2302399. PMID 37224232. Cited by the authors as “the largest comparative effectiveness trial to date”: ketamine was non-inferior to electroconvulsive therapy, in a mostly outpatient population. This result tempers the first place obtained here by electroconvulsive therapy, and the authors flag it themselves.

What was consulted. Verification carried out on 13 August 2026 against the full text of the publication, the open access PubMed Central version (33 pages), read in full: abstract, introduction, materials and methods, results, Figures 1 to 4 and their legends, discussion, acknowledgements, author contributions, funding, data availability, competing interests and the list of 30 references. All the figures published here, numbers of participants, odds ratios, confidence intervals, heterogeneity statistics, P-scores and response rates in the control arms, come from that text. The bibliographic metadata (journal, year, volume, issue, pagination, digital object identifier, PubMed identifier) were confirmed on the publisher’s website and on PubMed. The supplementary material (a 15.4 MB document and the PRISMA checklist) could not be consulted: it is hosted by the publisher and does not appear in the document consulted. Consequently, the supplementary figures and tables are not cited in this article, and no value that would exist only in that material is reproduced here. Two points found in the source itself are flagged to the reader: the breakdown of risk of bias differs between the abstract and the results section, and the PROSPERO registration number differs between the abstract and the methods section. The French regulatory status of the treatments cited was checked against the marketing authorisations in force, outside the source publication. Article subject to an independent double reading.

Editorial collections

Topics

Verified on 13 August 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on 16 September 2026, against the figures of the French version and against the source. How we verify what we publish.
Content published by Psychiatry Evidence Base is produced according to the principles of evidence-based medicine. Every analysis rests on an independent critical reading of the scientific literature and aims to help health professionals interpret it. The information presented replaces neither official guidelines, nor clinical reasoning, nor individualised care. Medicine evolves continuously, and some data may change as new scientific evidence appears.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigour.

Report an error in this analysis

Follow Dr Stroescu on LinkedIn, for the review every Saturday