Published on 20 September 2026
Brexpiprazole add-on for minimal and partial antidepressant responders: significant, but by a narrow margin
International Journal of Neuropsychopharmacology · 2025;28(10):pyaf074 · Kapadia et al.
DOI 10.1093/ijnp/pyaf074
PMID 41055581
Scientific 60
Editorial 79
The essentials
Three phase 3 trials comparing brexpiprazole 2 or 3 mg per day with placebo, added on top of an antidepressant, were pooled and reanalyzed after the fact according to how much patients had already improved during the eight weeks of antidepressant treatment preceding randomization. Among so-called minimal responders (improvement above 0% and below 25%, n=663), the difference in total MADRS score at six weeks reached −2.47 points (95% CI −3.38 to −1.55; p < 0.001; effect size 0.41). Among partial responders (improvement of 25% to under 50%, n=235), it was −1.53 points (−2.94 to −0.11; p = 0.035; effect size 0.28). Clinician-rated global severity moved in the same direction in both subgroups. What decides the reading is not the p-value: it is the size of the gap, about two and a half points on a scale that runs from 0 to 60, for a baseline score near 26. Add to this a non-pre-specified analysis whose plan was not registered, subgroups defined after the data were collected, no correction for multiple comparisons, no formal test of interaction between the two subgroups, and funding from the two companies that market the molecule and employ four of the six authors. Tolerability, at least, is reported in full, and it argues for caution: akathisia, motor restlessness and weight gain were markedly more frequent under brexpiprazole. This work produces a plausible hypothesis. It does not establish a clinical benefit.
Context
The situation is common in the consulting room. A patient started an antidepressant, the dose is right, so is the duration, and six or eight weeks later they feel somewhat better without feeling well. Neither a clear failure nor a clear response. The options are familiar and none of them is obviously the right one: raise the dose, switch molecules, add psychotherapy, or add a second agent, often a low-dose second-generation antipsychotic.
The trials that led to brexpiprazole being used as an add-on enrolled patients with an inadequate response to an antidepressant, broadly defined as less than 50% symptom reduction, without finely distinguishing degrees of that inadequacy. The question asked here is a fair one, and the authors state it plainly: clinical tradition tends to reserve switching antidepressants for patients who have barely moved, and augmentation for those who have partially responded. Within this heterogeneous group, is the benefit distributed evenly? A patient who has barely moved and a patient who has already improved substantially without remitting are not in the same clinical situation, and there is no reason to assume that adding a molecule gives them the same thing.
So the question is a good one. It is the way of answering it that calls for caution, and the format of a reanalysis conducted after the trials had ended cannot answer it in a confirmatory way. The authors say so themselves: the original trials were neither designed nor powered for this question, the analysis plan was not registered, and the results need to be confirmed by a pre-specified analysis.
What this analysis measures, and what it leaves out
It measures a change in score on a clinician-rated depression scale, the MADRS, at six weeks, using a mixed model for repeated measures. The MADRS was the pre-defined primary endpoint of each of the three trials. It also measures clinician-rated global severity, the CGI-S, a secondary endpoint, and the incidence of treatment-emergent adverse events. These are short-term efficacy and safety endpoints accepted in depression research, but they remain scale-based and event-reporting endpoints.
It does not measure functioning, quality of life, whether the benefit is maintained beyond six weeks, or what happens to patients after treatment stops. Nor does it compare brexpiprazole augmentation with the other strategies available in this situation, switching antidepressants or adding psychotherapy: no active comparator was planned, and the authors count this among their limitations. Finally, it reports neither response rates nor remission rates: a reader looking for how many patients crossed a clinically meaningful threshold will not find the answer here, only average changes.
The study at a glance
| Population | |
| DetailOutpatients aged 18 to 65, with a major depressive episode by DSM-IV-TR criteria lasting at least eight weeks, who had already received one to three antidepressants at adequate dose and duration with less than 50% improvement, followed by eight further weeks of a new, open-label antidepressant, that is, an inadequate response to two to four antidepressants in total. Any psychotic symptoms and other Axis I or II disorders specified in the protocol were exclusion criteria. Pooled efficacy sample: 1,162 patients, of whom 264 (22.7%) showed no improvement during the prospective period and were excluded from the analysis, 663 (57.1%) minimal responders and 235 (20.2%) partial responders. | |
| Intervention | |
| DetailBrexpiprazole 2 or 3 mg per day added to the ongoing antidepressant, for six weeks, after titration starting at 0.5 mg per day and increased to 1 mg per day after one week. Two trials randomized patients to 2 mg or placebo; the third randomized to 1 mg, 3 mg or placebo, and the 1 mg arm is not included in this analysis. An analysis restricted to the 2 mg dose appears in the supplementary material. | |
| Comparator | |
| DetailPlacebo added to the same antidepressant. | |
| Outcomes | |
| DetailChange in total MADRS score at six weeks, the pre-defined primary endpoint of each trial. Change in CGI-S, a secondary endpoint. Incidence of treatment-emergent adverse events, including those leading to discontinuation and those of an extrapyramidal nature. No response or remission rate is reported. | |
| Design | |
| DetailPooled post hoc analysis of three international, randomized, double-blind, placebo-controlled, fixed-dose phase 3 trials, conducted from 2011 to 2016 in Europe, Russia and North America (NCT01360645, NCT01360632, NCT02196506). Mixed model for repeated measures. Non-pre-specified analysis, analysis plan not registered. Oxford CEBM level 2b. |
Quality control
| Point checked | Judgment |
|---|---|
| Baseline data | Randomized, double-blind trials |
| FindingThe underlying material is solid: three phase 3, placebo-controlled trials, with blinding maintained on both arms during the six randomized weeks. The eight preceding weeks, however, were open-label for the investigator, with placebo masked only for the patient. The weakness is not in the data collection, it is in what was done with the data afterward. | |
| Sample size | Adequate for the main subgroup |
| FindingOf 1,162 pooled patients, 663 are minimal responders, which allows a reasonably precise estimate in that subgroup. The partial-responder group, n=235, is markedly smaller. Completion of the six weeks was high in both groups, 94.5% and 95.5% among minimal responders, 91.3% and 96.7% among partial responders. | |
| Statistical method | Model suited to missing data |
| FindingThe mixed model for repeated measures is the expected choice for this kind of follow-up, with treatment, visit, site nested within trial, baseline value and the corresponding interactions as factors. Confidence intervals for the between-group differences are provided, which allows precision, not just significance, to be judged. | |
| Analysis status | Post hoc, not pre-specified |
| FindingThe subgroups were defined after the data were collected, and the analysis plan was not registered. Baseline response was not a pre-planned stratification factor. This design changes the nature of what is produced: a hypothesis, not a demonstration. The authors themselves write that their results need to be confirmed in a pre-specified analysis. | |
| Multiple comparisons | No correction reported |
| FindingComparisons are tested at a nominal two-sided threshold of 0.05, with no adjustment for multiplicity, which the authors themselves flag as a factor inflating the risk of a type I error. Two efficacy endpoints, six weekly visits, two subgroups, plus the analyses in the supplement: the partial responders’ p=0.035 is the most exposed. | |
| Independence | Funder is also employer |
| FindingThe analysis is funded by Otsuka Pharmaceutical Development & Commercialization and Lundbeck, which market the molecule. Four of the six authors are or were employed by one or the other, and the remaining two declare paid consulting relationships, including with these same companies. The publication states that the sponsors took part in designing the research, in analyzing and interpreting the data, and in writing and reviewing the article. Writing assistance was also funded by them. | |
Results
| Result | Value reported |
|---|---|
| MADRS at six weeks, minimal responders (n=663) | −2.47 [−3.38; −1.55], p < 0.001, effect size 0.41 |
| FindingThe mean least-squares change was −8.8 points under brexpiprazole versus −6.3 points under placebo, with a standard error of 0.3 in both arms. The difference is statistically clear, and the confidence interval clearly excludes no effect. Superiority is present from week 1 and at every visit thereafter. Still, the gap between the two arms amounts to less than three points on a scale graded to 60, for a baseline score of 26.5. | |
| MADRS at six weeks, partial responders (n=235) | −1.53 [−2.94; −0.11], p = 0.035, effect size 0.28 |
| FindingMean changes of −6.4 versus −4.9 points, standard error 0.5 in both arms. The interval nearly touches zero at its upper bound and the p-value sits close to the conventional threshold. Without correction for multiple comparisons, this result is fragile. It should be neither taken as established nor dismissed as null. | |
| CGI-S at six weeks, minimal responders | −0.25 [−0.36; −0.13], p < 0.001, effect size 0.33 |
| FindingChanges of −1.1 versus −0.8 point, standard error 0.0 in both arms. The authors describe this result as consistent with the primary endpoint. On a scale with only seven grades, a quarter-point mean difference between groups remains a small change: the direction is genuinely consistent, the magnitude is not. | |
| CGI-S at six weeks, partial responders | −0.21 [−0.40; −0.01], p = 0.038, effect size 0.27 |
| FindingChanges of −1.0 versus −0.8 point, standard error 0.1. As with the MADRS in this subgroup, the upper bound of the interval nearly touches zero and the p-value falls just under the threshold. The result exists; it cannot carry much weight. | |
| Adverse events, overall incidence | 59.8% versus 47.8%; 54.8% versus 40.8% |
| FindingAmong minimal responders, at least one adverse event occurred in 196 of 328 patients under brexpiprazole versus 160 of 335 under placebo. Among partial responders, 63 of 115 versus 49 of 120. Discontinuations for an adverse event remained few, 2.7% versus 0.9% and 1.7% versus none. The gap in overall incidence, on the order of twelve to fourteen points, is wider than the efficacy gap observed on the rating scales. | |
| Adverse events of interest | Akathisia, motor restlessness, weight gain |
| FindingAmong minimal responders: extrapyramidal-type effects 14.3% versus 5.7%, akathisia 10.1% versus 3.9%, weight gain 5.5% versus 2.1%, motor restlessness 5.5% versus 0.6%. Among partial responders: extrapyramidal effects 13.9% versus 1.7%, akathisia 8.7% versus no cases under placebo, weight gain 6.1% versus no cases, headache 6.1% versus 2.5%, nasopharyngitis 5.2% versus 2.5%. These incidences are consistent with what had already been described for the full sample. | |
One detail is worth noting without over-interpreting it: the gap is larger among those who had improved the least than among those who had already partially responded. The authors offer an explanation, a higher baseline score in the first subgroup, 26.5 to 26.6 points versus 21.1 to 21.6, which mechanically leaves more room for improvement. Other explanations remain compatible with the observation, regression to the mean among patients selected for a weak response, or simply a difference in statistical power between two very unevenly sized subgroups. Above all, no test of interaction between the two subgroups is reported: nothing in the publication allows one to state that the difference between them is anything more than a fluctuation. Nothing, therefore, justifies turning it into an argument for clinical targeting.
Critical appraisal
| Domain | Judgment |
|---|---|
| Randomization and blinding | Preserved by the original design |
| FindingThe comparisons rest on groups randomized within each trial. The later split into subgroups is based on a variable measured before randomization, the degree of response during the prospective period, which limits, without eliminating, the risk of breaking comparability. The baseline characteristics reported are indeed close between arms within each subgroup. | |
| Analysis selection | Subgroups defined after the fact |
| FindingThis is the structuring limitation. When the cut points are chosen with knowledge of the data and the analysis plan is not registered, the reader cannot know how many splits were explored before the one that got published. The reported result should therefore be treated as hypothesis-generating. | |
| Alpha inflation | No adjustment reported |
| FindingTwo efficacy endpoints, six weekly visits, two subgroups, and no correction. The minimal responders’ result, with p < 0.001, would probably remain significant after adjustment. Those of the partial responders, with p = 0.035 on the MADRS and p = 0.038 on the CGI-S, are directly exposed. | |
| Effect size and clinical relevance | Small gap on a wide scale |
| FindingThe two gaps, 2.47 and 1.53 MADRS points, correspond to effect sizes of 0.41 and 0.28, which conventional benchmarks place between small and moderate. Two clarifications matter. The publication does not weigh its result against any threshold of clinical relevance, and proposes none: the assessment of magnitude offered here is an editorial reading, not a measure reported by the authors. And the values proposed in the literature for the smallest detectable difference on the MADRS vary with the estimation method, which rules out pitting this result against a single figure. The cautious conclusion is that a clearly perceptible benefit in the consulting room is not demonstrated by this data, not that it is excluded. | |
| Between-subgroup comparison | No interaction test |
| FindingThe article compares two estimates obtained separately, without formally testing their difference, and the trials were not powered for that comparison. With n=235, the partial-responder subgroup lacks power. Presenting the gap between the two subgroups as a finding would be a reading error, the mirror image of overselling the positive result. | |
| Independence and allegiance bias | Funder and employer are one and the same |
| FindingThe fact deserves stating without insinuation. The two companies that market the molecule fund the analysis, employ four of the six authors, and took part, by the publication’s own account, in the design, the analysis, the interpretation and the writing. The two remaining signatories declare paid consulting relationships. None of this invalidates a result. It does, however, constitute an allegiance risk documented in the methodological literature, one that bears less on the figures themselves than on the choices made upstream: which split to explore, which to publish, how to word the conclusion. This is precisely the kind of risk that a post hoc analysis, by construction, cannot neutralize. | |
| External validity | Trial population, six weeks |
| FindingThe patients enrolled are outpatients aged 18 to 65, free of psychotic symptoms or psychiatric comorbidity as specified in the protocol, first selected then further filtered by eight weeks of open-label antidepressant treatment. They are therefore less comorbid and more adherent than patients in routine practice, and the authors acknowledge this limit to generalizability. The trials were conducted from 2011 to 2016 in Europe, Russia and North America. The six-week horizon says nothing about whether the benefit is maintained, or about what happens after treatment stops. For lack of a sufficient number of men, no analysis by sex could be conducted. | |
| Safety | Excess akathisia and weight gain |
| FindingUnlike many post hoc efficacy analyses, this one reports its tolerability data, and they are not neutral. In both subgroups, extrapyramidal-type effects were markedly more frequent under brexpiprazole than under placebo, 14.3% versus 5.7% among minimal responders and 13.9% versus 1.7% among partial responders, and the same holds for akathisia and for weight gain. Follow-up lasts only six weeks, which necessarily underestimates metabolic consequences: the long-term analyses cited by the authors report an average weight gain on the order of four kilograms over one year. When the expected gain on the rating scale is two to three points, these elements weigh heavily in the balance. | |
Level of evidence
Confidence is high on one point, and one point only: in these three pooled trials, among patients whose improvement under an antidepressant had been the weakest, adding brexpiprazole 2 or 3 mg was accompanied by a larger drop in MADRS than placebo at six weeks, and this difference is not plausibly attributable to chance. The consistency between the MADRS and the CGI-S, which move in the same direction with similar effect sizes, reinforces this finding.
Confidence is low everywhere else. On whether a real difference exists between minimal and partial responders, which here rests on an unplanned comparison, never formally tested, between two subgroups of unequal power. On the clinical translation of the gap, whose magnitude is small relative to the scale’s range and to the baseline score. On reproducibility, which a non-pre-specified analysis with no correction for multiple comparisons cannot anticipate, and which the authors themselves call to confirm in a pre-specified analysis. On the independence of the work, finally. What is demonstrated fits in one sentence: in this dataset, a gap exists, and it comes with a measured excess of adverse events. What is suggested fits in another: this gap could be larger among patients who respond the least. And the view that this gain is not enough to change a treatment strategy is, at this stage, reasoned opinion rather than proof.
The colleague test
What an experienced colleague would say if handed this study for two minutes, between two consultations.
“My patient barely moves on his antidepressant, so yes, the question interests me. And the signal is there, p below 0.001, that’s not nothing. Only, two and a half points of MADRS, I’m not sure I’d know how to spot that in a consultation, and meanwhile akathisia goes from 3.9% under placebo to 10.1% under brexpiprazole. The analysis was done after the fact, and it’s the company that funds it and employs most of the signatories. I’m not changing my practice over this. I go back to the pivotal trials, and I look at the other options first.”
Translation for practice: this analysis provides no new argument for widening the use of brexpiprazole augmentation. It does, however, provide a good reading exercise, on the distance between a significant p-value and a benefit the patient actually feels.
A note on regulatory status
This point falls outside the study itself, but it conditions how the result can be used, and it is not trivial: authorization, actual marketing and collective reimbursement are three separate questions, and a decision on any one of them says nothing about the others, least of all about demonstrated clinical efficacy. A negative reimbursement opinion, in particular, never means a treatment is ineffective. France offers a verified, dated example of how these three planes can diverge for the same molecule; the elements below were checked on August 13, 2026, with the European Medicines Agency and the Haute Autorité de santé, and readers should verify the corresponding texts in their own jurisdiction, since they evolve.
In the European Union, brexpiprazole is authorized under the name Rxulti since July 26, 2018, held by Otsuka Pharmaceutical Netherlands B.V., and its marketing authorization is limited to the treatment of schizophrenia in adults and adolescents aged 13 and over. Antidepressant augmentation in major depressive episode is not among the authorized indications. In the United States, by contrast, this adjunctive use in adults has been approved since July 2015: this asymmetry explains why the three trials pooled here fed an indication on one side of the Atlantic and not on the other.
In France, the Haute Autorité de santé’s Commission de la Transparence issued an opinion on November 19, 2025 on Rxulti 1, 2, 3 and 4 mg in its schizophrenia indication: insufficient clinical benefit (service médical rendu) to justify reimbursement by national solidarity given the alternatives available, and an unfavorable opinion on listing. Authorization, actual marketing and reimbursement remain three distinct questions, and the state of availability in any given country should be verified at the source before prescribing, since these texts change.
It follows that, in France, prescribing brexpiprazole as an antidepressant augmentation would be an off-label use. Off-label prescribing is not freely permitted everywhere: in France, Article L. 5121-12-1 of the Code de la santé publique, in the version in force as of December 28, 2023, makes it conditional on the absence of an appropriate authorized treatment, specific information given to the patient about expected risks and benefits, a notation on the prescription, and a justification recorded in the medical file. The conditions governing off-label prescribing vary from one country to another, and clinicians should check the framework of their own jurisdiction before relying on this option.
What you can do with this.
- Do not change a strategy on the basis of this reanalysis alone. The pivotal trials and current guidelines remain the reference for deciding on an augmentation.
- Before considering adding a second agent in a patient with a poor response, revisit the basics: accuracy of the diagnosis, actual dose and duration of the antidepressant, real adherence, somatic comorbidity, alcohol or substance use, untreated sleep disorder.
- If an augmentation is chosen, frame it from the outset as a bounded trial: a defined duration, a response criterion measured with a scale rather than an impression, and a decision to stop if the expected gain does not appear.
- Monitor what this analysis actually measured on the risk side: akathisia and motor restlessness, whose excess is clear from the first weeks, and weight, whose course cannot be judged over six weeks.
- Keep this study as a teaching tool. It illustrates well, with a resident or in a peer group, the difference between a statistically significant result and a result that changes something for the patient.
- The course of action is set out in the NICE decision tree for schizophrenia in adults.
Frequently asked questions
Can I conclude that adding brexpiprazole helps patients who barely respond at all?
No, not from this work alone. It shows that a gap exists in a pooled dataset, about two and a half MADRS points on a scale that runs to 60, with an effect size of 0.41. The publication reports neither a response rate nor a remission rate, so nothing here says how many patients crossed a clinically meaningful threshold. A real but small average difference is not, on its own, a therapeutic argument.
Why does the effect look larger in patients who were doing worst at baseline?
The analysis cannot answer that. The authors point to the higher baseline score in this subgroup, which mechanically leaves more room for improvement. Regression to the mean in patients selected for a weak response, or simply a difference in statistical power between a subgroup of 663 and one of 235, remain equally compatible with the observation. No test of interaction between the two subgroups is reported, which rules out claiming that the difference between them is real.
Does the weaker result in partial responders mean the strategy doesn’t work in this group?
No, and the result is not negative either: the difference reaches the conventional threshold, on both the MADRS and the CGI-S, but narrowly and without correction for multiple comparisons. With n=235, precision is limited in both directions. Failing to find a larger effect is not the same as finding no effect, and narrowly crossing a threshold is not the same as establishing one.
Is industry funding enough to dismiss these results?
No, and saying so would be as careless as ignoring it. The data come from randomized, double-blind trials, and the source of funding does not make a number false. What it makes hard to rule out is influence over decisions that stay invisible in the article: which splits were explored, which were kept, how the conclusion was worded. The publication itself states that the sponsors took part in the analysis, the interpretation and the writing. This is exactly where a non-pre-specified analysis leaves the most room, which is why the two limitations reinforce each other.
What do we know about tolerability in this analysis?
It is reported, and it deserves to be read before the efficacy findings. At least one adverse event occurred in 59.8% of patients under brexpiprazole versus 47.8% under placebo among minimal responders, 54.8% versus 40.8% among partial responders. Extrapyramidal-type effects, akathisia and weight gain are what widen the gap. Discontinuations for an adverse event remained rare over six weeks, but six weeks says nothing about the metabolic side.
Annotated bibliography
Kapadia S, Zhang Z, Ardic F, Patel M, Thase ME, Papakostas GI. Adjunctive brexpiprazole in patients with major depressive disorder who show minimal or partial response to antidepressant treatment: post hoc analysis of randomized controlled trials. International Journal of Neuropsychopharmacology. 2025;28(10):pyaf074. DOI 10.1093/ijnp/pyaf074. PMID 41055581. Pooled post hoc analysis of three randomized phase 3 trials comparing brexpiprazole 2 or 3 mg per day with placebo as an antidepressant augmentation, in adults with a major depressive episode. What it contributes: an estimate of the gap from placebo, on the MADRS and the CGI-S, in two subgroups defined by the degree of response achieved during the eight weeks of antidepressant treatment preceding randomization, together with adverse event incidences. Its limitations: a non-pre-specified analysis whose plan was not registered, no correction for multiple comparisons, no test of interaction between subgroups, no active comparator, a six-week duration, and funding and affiliation of four of the six authors with the companies that market the molecule.
Original trials pooled in this analysis, as identified in the publication: NCT01360645, NCT01360632 and NCT02196506 (ClinicalTrials.gov), conducted from 2011 to 2016 in Europe, Russia and North America.
What was checked. The article underwent an independent double reading. The figures, sample sizes, confidence intervals, p-values, effect sizes and adverse event incidences were checked one by one against the full text of the publication, in its open-access version deposited on PubMed Central, and against the publisher’s record for the title, authors, volume, pagination and persistent identifiers. The publication’s supplementary material, which contains the analysis restricted to the 2 mg per day dose, could not be consulted: no data drawn from it appears in this article. The regulatory elements were verified on August 13, 2026 with the European Medicines Agency, the Haute Autorité de santé and Légifrance; these sources evolve and should be re-checked before any decision-making use.
Editorial collections
Tags
Verified on August 13, 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on September 20, 2026, against the figures of the French version and against the source. How we verify what we publish
This analysis is intended for healthcare professionals. It does not constitute a prescribing recommendation and does not replace individual clinical judgment.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigor.
