Published on 15 September 2026
Esketamine in treatment-resistant depression: significant, and below the trials’ own threshold
In brief
Seven randomised, double-blind, placebo-controlled trials, six of initiation and one of continuation, 1,505 patients, and above all the individual participant data obtained from the manufacturer through the Yale Open Data Access Project. An eighth trial, in monotherapy, was analysed separately on aggregate data, its individual data not having been supplied despite the request. An independent team reanalysed the material that served to license intranasal esketamine in treatment-resistant depression, under a protocol the journal accepted before any result was known. In the five initiation trials combining esketamine with an antidepressant that document this outcome, the mean difference on the MADRS at four weeks or more is −2.94 points, 95% CI −5.39 to −0.48, moderate GRADE certainty. The effect exists, it points the right way, and the whole interval stays below the 6.5-point threshold that the pivotal trial protocols had themselves used to size their samples. Sedation, dissociation and adverse events are markedly more frequent. No excess of serious events is demonstrated, but the incidence rate ratio is 1.35 with an interval of 0.54 to 3.40: a tripled risk remains compatible with these data, on a sample that cannot rule it out. The authors conclude to an advantage that is statistically significant but small, and whose clinical relevance remains unclear given the risks.
The context
Intranasal esketamine was approved in 2019 by the FDA and by the EMA in treatment-resistant depression, on a dossier made of trials run by the manufacturer. Since then the drug has held a particular place: awaited by patients who have run out of options, framed by a heavy organisation, and discussed in the methodological literature almost as much as in the clinical literature. The question asked here is not whether esketamine does something. It is how much, measured on the raw data rather than on the papers that report them.
Two features of method change the status of this work. The first is the format: a registered report, whose protocol was posted on the Open Science Framework and granted in-principle acceptance by the journal on 15 November 2021, then registered on PROSPERO on 11 December 2021 under the number CRD42021290721. The second is the source: individual participant data, not the aggregate results of the publications alone. The chronology is the main argument here. The data access requests were submitted to the YODA platform on 5 January 2022, and access was granted on 8 February 2022. The analysis plan was therefore fixed and approved by the journal before the authors had even asked for the data, which removes most of the room for manoeuvre after the fact.
One precedent shows what the approach is worth. After the alerts published by this same team on the continuation trial, the sponsor reported a sensitivity analysis whose p value moved from a highly significant result to one described as “barely significant”. That is exactly the kind of gap that stays invisible in the primary publications and that access to individual participant data allows one to examine.
Then there is the benchmark, which is the heart of the paper. The 6.5-point threshold on the MADRS is not a fact of nature, and it was not invented by the authors of the reanalysis. On their account it comes from the protocols of the three first pivotal trials, where it served the sample size calculation, on the basis of phase 2 results and of a “clinical judgement”. In other words, it is the step the development programme had set for itself. It is neither consensual nor validated elsewhere, but it has the rare advantage of having been written beforehand, and by the party the comparison works against.
The study at a glance
| Question (PICO) | |
|---|---|
| Population | |
| Adults with treatment-resistant depression as defined by the trials of the esketamine development programme. 1,505 patients, individual participant data from 7 randomised, double-blind, placebo-controlled trials, that is 6 initiation trials and 1 continuation trial. An 8th trial, in monotherapy (476 participants), was analysed separately on aggregate data, its individual data not having been supplied despite the request. Most patients had a low level of resistance | |
| Intervention | |
| Intranasal esketamine, combined with an oral antidepressant in the initiation trials and in the continuation trial, or as monotherapy in the 8th trial analysed separately | |
| Comparator | |
| Intranasal placebo, with or without an antidepressant depending on the trial | |
| Primary outcome | |
| MADRS score at 4 weeks or more in the initiation trials, benchmarked against the clinical significance threshold of 6.5 points used in the design of the pivotal trials. Certainty rated with GRADE | |
| Design | |
| Systematic review and individual participant data meta-analysis, published as a registered report. Two-step meta-analysis for efficacy, one-stage for the search for moderators. Separate analyses by phase (initiation or continuation) and by treatment type (combination or monotherapy) · CEBM 1a · PROSPERO CRD42021290721 |
Quality control
| Criterion | Status |
|---|---|
| Pre-registration | Sound |
| FindingRegistered report granted in-principle acceptance before any result was known, protocol public on the Open Science Framework, PROSPERO registration | |
| Nature of the data | Sound |
| FindingIndividual participant data obtained through the Yale Open Data Access Project, not a rereading of the published results alone | |
| Selection and risk of bias | Sound |
| FindingTwo independent researchers for study selection and for the risk of bias assessment | |
| Choice of the benchmark threshold | Sound |
| FindingThe 6.5-point threshold is the one in the protocols of the pivotal trials, so it predates this reanalysis and was not picked after the fact | |
| Access to centre-level data | Reservation |
| FindingThe centre identifier had been removed from the datasets supplied, for anonymisation reasons. No centre-by-centre analysis was therefore possible, and the authors proceeded by excluding countries | |
| Data integrity | Reservation |
| FindingThe FDA appraisal describes possible problems in two trials. For the initiation trial conducted in elderly participants: an unusual response curve shift, discrepancies between locked datasets and reported protocol violations. For the continuation trial: a site in Poland where the entire placebo arm relapses and which drives the overall result on its own | |
| Monotherapy trial | Reservation |
| FindingAggregate data only, with reservations from the authors about a selection bias and about possible unblinding | |
| Funding and independence | Sound |
| FindingFunded by a hospital clinical research programme of the French Ministry of Health, with no role of the funder in the design, the data collection, the analysis, the decision to publish or the preparation of the manuscript. Three authors are members of a list of industry-independent experts, and none declares any interest with the manufacturers of the drugs under study | |
The findings
| Outcome | Result |
|---|---|
| MADRS at 4 weeks or more, combination (5 trials) | MD = −2.94 (95% CI −5.39 to −0.48), moderate GRADE certainty |
| PEB readingSignificant but the whole interval stays below the 6.5-point threshold | |
| Standardised effect size | SMD = −0.25 (95% CI −0.45 to −0.06) |
| PEB readingSmall effect Value taken from the full text, absent from the published abstract | |
| MADRS at 1 week (6 initiation trials) | MD = −2.51 (95% CI −4.81 to −0.20) |
| PEB readingSame order of magnitude from the first week Value taken from the full text | |
| Relapse, continuation trial | HR = 0.38 (95% CI 0.26 to 0.57) |
| PEB readingClear reduction of the risk but the primary outcome loses its significance once all the Polish centres are excluded | |
| Monotherapy, 1 trial, aggregate data | MD = −6.32 (95% CI −8.62 to −4.03) |
| PEB readingLarger effect but explicit reservations about selection bias and about unblinding | |
| Sedation | RR = 3.70 (95% CI 2.02 to 6.78) |
| PEB readingMultiplied by nearly 4 | |
| Dissociation | RR = 2.36 (95% CI 2.10 to 2.65) |
| PEB readingNarrow interval, a constant and expected adverse effect | |
| Adverse events | IRR = 3.91 (95% CI 2.37 to 6.45) |
| PEB readingClear excess | |
| Serious adverse events | IRR = 1.35 (95% CI 0.54 to 3.40) |
| PEB readingNo excess demonstrated but the interval leaves room for a more than tripled risk | |
| Moderators | 5 moderators tested, all non-significant: resistance level p = 0.80, age p = 0.45, baseline severity p = 0.79, gender p = 0.87, class of the combined antidepressant, SSRI or SNRI, p = 0.91 |
| PEB readingNo better-responding patient profile emerges Most patients had low-stage resistance, which limits what can be said about the most resistant forms | |
The figure to hold on to is not an absolute one, it is a relative one. A difference of 2.94 points on a scale that runs to 60 corresponds to a standardised effect size of 0.25. That is real and it is small. The upper bound of the interval, at −5.39 points, still falls short of 6.5. So this is not a power problem: even at the most favourable bound of the confidence interval, the step the development programme had set for itself is not reached.
On relapse prevention, the hazard ratio of 0.38 is the most impressive result of the set, and it is also the most fragile. The centre identifier had been removed from the datasets supplied for anonymisation reasons, which ruled out any analysis excluding a given centre. The authors therefore proceeded by excluding countries. When all the Polish centres are removed, the effect size falls and the primary outcome, remission, loses its statistical significance. The response outcome keeps it. That detail does not appear in the published abstract and reads as a signal of fragility, not as a refutation.
Three elements complete the reading. First, the FDA appraisal described possible data integrity problems in two trials: an unusual response curve shift, discrepancies between locked datasets and reported protocol violations for the initiation trial conducted in elderly participants, and a site in Poland whose entire placebo arm relapses for the continuation trial. Second, the search for moderators is negative on the five variables tested, gender included with a p of 0.87, which leaves no empirical basis for designating in advance a patient profile that would benefit more. Third, the work is funded by a hospital clinical research programme of the French Ministry of Health, the funder having played no part in the design, in the data collection and analysis, in the decision to publish, or in the preparation of the manuscript.
Critical appraisal
| Domain | Judgement |
|---|---|
| Protection against opportunistic analyses | Sound |
| FindingRegistered report, analysis plan fixed before the results, protocol accessible. The format does not guarantee that the answer is right, it guarantees that the answer was not chosen | |
| Fit between claim and evidence | Sound |
| FindingThe authors’ conclusion, an effect that is statistically significant but small and of unclear clinical relevance, matches exactly what the intervals show | |
| Validity of the benchmark threshold | Reservation |
| FindingThe 6.5-point threshold comes from the industry protocols and from a clinical judgement, not from a study of the minimal clinically important difference. It is legitimate as an internal benchmark, it does not have the standing of a universal norm | |
| Robustness of the relapse result | Reservation |
| FindingA result carried by a single continuation trial. Anonymisation ruled out any centre-by-centre analysis, and excluding all the Polish centres makes the primary outcome lose its significance | |
| Integrity of blinding in the source trials | Reservation |
| FindingWith a relative risk of dissociation of 2.36, the blind is hard to keep watertight. This reasoning is a PEB reading, not a result of the paper, but it weighs on the interpretation of an effect of this size | |
| External validity towards the most resistant forms | Reservation |
| FindingThe population analysed was mostly of low resistance. The absence of moderation by resistance level cannot be extrapolated to the heaviest patients, who are nevertheless the ones the drug is offered to | |
| Duration of observation | Reservation |
| FindingThe primary efficacy outcome is measured at four weeks or more. Only the continuation trial documents the maintenance of the benefit, and that is besides the most fragile result of the dossier. Tolerability over several months of real-world use is not covered by this material | |
| Orientation of the authors | Reservation |
| FindingThe team includes researchers known for their critical work on the evaluation of antidepressants and on publication bias. The registered report format limits what that orientation could do to the results, it does not settle it as a question | |
One distinction deserves to be held firmly. This reanalysis does not show that esketamine is ineffective: the direction of the effect is consistent, it appears from the first week, and the interval excludes zero. It shows that the size of the effect, measured on the data of the registration dossier, stays below the step of clinical relevance written into the protocols of that same dossier. These are two different statements, and the second is the only one the data allow.
Level of evidence
PEB appraisal: high confidence on the existence and on the smallness of the effect in combination, moderate confidence on relapse prevention, low confidence on monotherapy. The methodological apparatus is the best available for reappraising a registration dossier: individual participant data, protocol accepted before the results, independent double assessment, explicit GRADE certainty. What lowers confidence has nothing to do with the conduct of the work and everything to do with the material: six short initiation trials, a single continuation trial, an eighth trial in monotherapy available as aggregate data only, and a population less resistant than the one the drug is offered to in practice.
The colleague test
What an experienced colleague would say if you put this study to them in two minutes, between two consultations.
“ Three MADRS points. They had set themselves six and a half to calibrate their own trials, and we do not get there, not even taking the most favourable bound. I will keep referring patients who have very little left, but I will stop letting people believe it is a turning point. And I say it before the first session, not when the patient asks me why he feels nothing. ”
What this means in practice: the indication does not move, the words that go with it do. We move from a promise of reversal to a measured proposal, with a modest average benefit and markedly more frequent adverse effects. No responder analysis is reported, and the five moderators tested are negative: nothing allows anyone, as things stand, to tell a given patient that he is one of those who will respond strongly.
What you can do with this on Monday morning. The indication stays the same. What changes is what you announce, to whom, and in what order.
- Put figures on the information before the first administration rather than after: an average gain of the order of 3 MADRS points at four weeks or more, a frequency of sedation multiplied by 3.7 and of dissociation by 2.4 compared with placebo. The source publishes no absolute frequencies, so these ratios do not translate directly into a percentage of sessions concerned.
- Do not reserve esketamine for the heaviest forms in the hope of an effect proportional to resistance: no moderation by resistance level was found. The reservation to state in the other direction is that the population analysed was of low resistance, so this absence of moderation has not been tested where you prescribe most.
- Go back over the precautions of the summary of product characteristics in force where you practise at every session and not only at initiation: administration under the supervision of a health professional, blood pressure measured before and then about 40 minutes after the dose, no driving and no operating machinery until the next day after a restful night’s sleep. Contraindications to recheck: hypersensitivity to the active substance, to ketamine, or to any of the excipients, aneurysmal vascular disease, including intracranial, thoracic or abdominal aortic, or peripheral arterial vessels, any history of intracerebral haemorrhage, with no criterion of how long ago, and a recent cardiovascular event, defined as less than six weeks old. The United States labelling adds arteriovenous malformation alongside aneurysmal vascular disease; the European labelling does not mention it. Labelling is decided jurisdiction by jurisdiction, so the wording that applies is the one in force where you practise.
- Document the decision to continue beyond the initiation phase. The reduction in relapse risk is the best argument in favour of continuation, and it is also the most fragile result of the dossier. A dated reassessment, with an explicit criterion, is worth more than a tacit renewal.
- Use this work as a teaching case with trainees: a threshold of clinical relevance written before the results, by the party the comparison works against, is worth every commentary on the difference between significance and relevance.
Frequently asked questions
Should esketamine no longer be prescribed?
No, and the paper does not suggest it. The effect is real, it points the right way, and the drug keeps its place in patients whose options can be counted on the fingers of one hand. What the reanalysis forces us to revise is the size announced, not the existence of the indication.
A result that is significant but below the threshold, what does that mean in practice?
That the difference between esketamine and placebo is probably not due to chance, and that it stays small. Statistical significance answers the question of existence, not that of size. Here the two answers diverge, and that is the whole point of the paper.
Where does the 6.5-point threshold come from?
From the protocols of the three first pivotal trials, where it served the sample size calculation, on the basis of phase 2 results and of a clinical judgement, as the authors describe it. It is not a validated norm for the minimal clinically important difference, and it should not be presented as one. Its strength lies elsewhere: it was written beforehand, by the development programme itself.
Do the most resistant patients get more out of it?
Nothing shows that. Five moderators were tested, resistance level, age, baseline severity, gender and class of the combined antidepressant, and none is significant. Be careful, though, not to read that absence as a demonstration: most of the patients included had low-stage resistance, which leaves the question open for the most severe forms.
Is a team that has been criticising esketamine for years credible to reappraise it?
The question is legitimate and the answer lies in the format. A registered report fixes the analysis plan before the results are known and has the journal approve it at that stage. The authors’ orientation could influence the choice of questions, far less the handling of the data. That is precisely what this format is meant to produce, and it is why it should be required more often, of industry too.
Annotated bibliography
Naudet F, Pellen C, Fodor LA, Gastaldon C, Barbui C, Turner EH, Le Pabic E, Cristea IA (2025). Efficacy and safety of esketamine for “treatment resistant depression”: registered report for a systematic review with an individual patient data meta-analysis of randomized, double-blind, placebo-controlled trials. BMC Medicine, 23(1), 677. DOI 10.1186/s12916-025-04435-x · PMID 41310599. Source study analysed here. PROSPERO registration CRD42021290721. Funded by a hospital clinical research programme of the French Ministry of Health (ESK-T-Dep), with no role of the funder in the design, the data collection, the analysis, the decision to publish or the preparation of the manuscript.
Popova V et al. (2019). Efficacy and Safety of Flexibly Dosed Esketamine Nasal Spray Combined With a Newly Initiated Oral Antidepressant in Treatment-Resistant Depression: A Randomized Double-Blind Active-Controlled Study. American Journal of Psychiatry. DOI 10.1176/appi.ajp.2019.19020172 · PMID 31109201. A pivotal initiation trial, one of those whose individual participant data were reanalysed here. Worth reading to see how the same material looks depending on whether it is reported or reanalysed. Note the title: both arms received a newly initiated oral antidepressant, the comparator of the nasal spray was a placebo, so this is not an active comparator of esketamine. Limitations: trial run by the manufacturer, limited sample.
Daly EJ et al. (2019). Efficacy of Esketamine Nasal Spray Plus Oral Antidepressant Treatment for Relapse Prevention in Patients With Treatment-Resistant Depression: A Randomized Clinical Trial. JAMA Psychiatry, 76(9), 893-903. DOI 10.1001/jamapsychiatry.2019.1189 · PMID 31166571. The continuation trial, source of the relapse hazard ratio of 0.38 carried into the meta-analysis. It is also the trial one of whose Polish sites drew reservations from the FDA. Anonymisation of the datasets ruled out a centre-by-centre analysis, hence the exclusion of countries adopted as a substitute approach.
Reif A et al. (2023). Esketamine Nasal Spray versus Quetiapine for Treatment-Resistant Depression. New England Journal of Medicine. DOI 10.1056/NEJMoa2304145 · PMID 37792613. An open-label trial comparing esketamine and extended-release quetiapine as add-on, favourable to esketamine. A useful counterpoint, with two limitations that rule out setting it against the reanalysis: an open design, so no blinding, and industry sponsorship.
