Published on 16 September 2026

Analysis · Bipolar disorder · Psychopharmacology

Manic switch on antidepressants in bipolar depression: what can a network meta-analysis settle?

▤ Dossier eClinicalMedicine · 2025; 87: 103413 · Oliva et al. DOI 10.1016/j.eclinm.2025.103413 PMID 40823496 Scientific 72 Editorial 87

In brief

The fear of tipping a patient with bipolar disorder into a switch by starting an antidepressant has shaped prescribing for decades. Network meta-analyses had already compared antidepressants with other drug classes, but none, according to the authors, had restricted itself to a network made up only of antidepressants and placebo on this outcome. That is what this work does: 13 randomised controlled trials, 1362 patients in an acute depressive episode of bipolar disorder, 12 antidepressants plus placebo. The main result is negative in the strict sense: no antidepressant stands out with a risk of manic or hypomanic switch significantly higher than placebo, and no head-to-head comparison reaches significance. This result has to be read with the certainty the authors themselves assign to it. The CINeMA assessment covers 78 comparisons: 77 are rated low confidence, the last one very low confidence. None escapes that reservation. Venlafaxine carries the only notable point signal, a risk ratio of 4.53 with a 95% confidence interval running from 0.47 to 43.25, and therefore compatible with no difference, but it is also the only compound for which the authors note consistent signals in individual trials. The P-score ranking puts amineptine and amitriptyline at the top, placebo in third place and venlafaxine last: this is a relative rank computed inside a network that carries little information, not a demonstrated safety hierarchy. Finally, mean trial duration is 6.9 weeks, with a maximum of 10 weeks. The article therefore does not demonstrate that antidepressants are safe in bipolar depression. It documents, with no apparent overinterpretation, how much these trials leave unresolved.

The context

The question asked here is among the most ordinary in the consulting room and among the most poorly supported. A patient known to have bipolar disorder comes in during a depressive phase, the suffering is real, the mood stabiliser alone has not been enough, and the question of adding an antidepressant arises. The answer depends on a trade-off between an antidepressant benefit that is itself debated in this indication and a risk of tipping into hypomania or mania of which every clinician has examples in mind.

The literature was not empty. The largest network meta-analysis devoted to the acute treatment of bipolar depression had already found more switches on antidepressants than on other classes, with no difference from placebo, and without any single drug standing out. What this work adds, according to its authors, is narrower: a network restricted to antidepressants and placebo alone, intended to improve the resolution of comparisons within the class.

The clinical issue is therefore not whether antidepressants cause switches, but which ones. Practice runs on empirical rules, including particular caution with venlafaxine, resting on clinical case series and international guidelines rather than on a quantitative comparison of randomised trials. The merit of this work is to put the question in the terms that would allow it to be answered. The reader should anticipate the shift that follows: the article does not close the debate on switching, it establishes the state of the evidence.

The methodological point

What the P-score says, and what it does not

A network meta-analysis compares treatments that have never been tested against each other, by going through their common comparators. It thus widens the number of comparisons available and, in return, inherits the weaknesses of the trials that feed it, with one additional assumption, transitivity: the populations and conditions of the trials must be close enough for the indirect reasoning to hold. The model used here is frequentist, with random effects, and not Bayesian.

The P-score is a statistic derived from this network. For each treatment it expresses the mean certainty of being safer than the others on the outcome considered, and it runs from 0 to 1. It always produces a ranking, whatever the quality of the information available. That is where the most common misreading takes place: a network in which no comparison reaches significance still produces a complete order, from first to last. The rank exists because the statistic computes it, not because the data demonstrate it. When the intervals overlap widely, the ranking essentially describes noise.

One detail makes this reading particularly instructive here: placebo sits in third place in the ranking, behind only two antidepressants. Ten of the twelve drugs therefore have a point estimate less favourable than placebo, without any of them reaching significance. From this the authors draw a cautious wording, “suggesting a possible, albeit non-significant, class effect”, and not the opposite conclusion.

The CINeMA tool corrects the illusion of the ranking by assigning each comparison in the network a level of confidence that takes into account the risk of bias of the trials, imprecision, heterogeneity, incoherence between direct and indirect evidence, and indirectness. That the authors applied CINeMA and published the result without softening it is a strength of this work. That the result is low for 77 comparisons and very low for the last one is what should govern the reading of the rest of the article, including the ranking.

That leaves the question of the outcome itself. Switch is defined by the authors of the original trials, with no additional operational criterion. Eight trials used standardised criteria or rating scales, with thresholds that do not coincide: a Young Mania Rating Scale score above 10, above 11, of 12 or more, of 13 or more, or of 16 or more depending on the trial, and elsewhere categorical DSM criteria, the Raskin scale or the DOTES scale. The remaining five trials relied on clinical observation. These definitions do not delimit the same event. A low threshold captures mild and transient mood elevations, a high threshold retains only full episodes. Pooling these definitions amounts to measuring different objects under one name, and the expected effect is dilution. The authors ran a sensitivity analysis on this point, which did not change the results, while acknowledging that “subtle inconsistencies may have gone undetected”.

The study at a glance

Question (PICO)
Population
1362 patients in an acute depressive episode of bipolar disorder, diagnosed according to various versions of the DSM. 688 patients, or 50.5%, had bipolar I disorder, 183 or 13.4% bipolar II disorder, and the subtype was not reported for 491 patients, or 36.1%. Women: 818, or 60.1%.
Intervention
12 antidepressants: agomelatine, amineptine, amitriptyline, bupropion, desipramine, fluoxetine, imipramine, moclobemide, paroxetine, sertraline, tranylcypromine, venlafaxine. The network is therefore dominated by older drugs, tricyclics and monoamine oxidase inhibitors included. 930 patients, or 68.3%, received the antidepressant as add-on to a mood stabiliser, 432 or 31.7% as monotherapy.
Comparator
Placebo or another antidepressant. Only six trials included a direct comparison with placebo, the other comparisons resting on the network.
Outcome
Occurrence of a switch to mania or hypomania after antidepressant treatment during the acute depressive episode, as reported by the authors of each trial, with no additional operational criterion. Intention-to-treat analysis in all included trials.
Design
Systematic review and frequentist random-effects network meta-analysis of 13 parallel-group randomised controlled trials, protocol preregistered on the Open Science Framework, reporting following PRISMA-NMA, risk of bias assessed with RoB 2, confidence in the estimates assessed with CINeMA. Search up to 19 February 2025, with no language restriction. The level of evidence 1a according to the Oxford CEBM classification is an editorial judgement, external to the publication.

The quality check

CriterionStatus
Protocol registrationSound
FindingProtocol preregistered on the Open Science Framework, with amendments reported in the appendix. The risk of analyses decided after looking at the data is limited, though not eliminated, since the authors themselves describe their post hoc analyses by mechanism of action as exploratory.
Assembly of the evidence baseSound
FindingThirteen parallel-group randomised trials, forming a connected network that allows indirect comparisons. Connectedness is a necessary condition for the analysis, it says nothing about the density of the network: only six trials compared an antidepressant directly with placebo.
Risk of bias of the trialsReservation
FindingAssessment with RoB 2: high risk in two trials, some concerns in four others, low in the rest. A sensitivity analysis excluding the two high-risk trials gives results consistent with the main analysis.
Definition of the outcomeReservation
FindingSwitch is defined by different thresholds and methods from one trial to another, from the Young Mania Rating Scale to clinical observation alone in five of them. Pooling events that do not overlap mechanically dilutes the differences between drugs.
Confidence in the estimatesLow confidence
FindingThe CINeMA assessment covers 78 comparisons: 77 are rated low confidence and one very low confidence. No comparison in the network therefore escapes that reservation.
Length of observationReservation
FindingMean trial duration of 6.9 weeks, ranging from 1.7 to 10 weeks. Switches occurring beyond this window are not measured here, which allows no conclusion that they do not occur.
Heterogeneity and consistencySatisfactory
FindingNull between-study variance and I² at 0%, but with a 95% confidence interval reaching 79.2%, which reflects how little information is available more than an established homogeneity. No global or local inconsistency detected, no publication bias identified.

The findings

78 / 78Comparisons in the network whose CINeMA confidence is low, 77 of them, or very low, the last one. This is the figure that conditions the reading of all the others, including the ranking of the drugs.

The central result is an absence of statistically significant difference: none of the 12 antidepressants assessed shows, compared with placebo, a risk of switch whose interval excludes no effect, and no head-to-head comparison reaches significance. This wording is heavier than saying that antidepressants do not cause switches, and it is the only one the data allow. In a network where every comparison carries low or very low confidence, a non-significant comparison settles practically nothing: it records that the information available is insufficient. The symmetry holds in both directions, since the same data do not allow anyone to claim either that a given antidepressant is safer than another.

Risk ratios against placebo range from 0.31 for amitriptyline to 4.53 for venlafaxine. Ten of the twelve drugs have a point estimate above 1. The post hoc pairwise meta-analysis restricted to the six trials comparing an antidepressant directly with placebo gives a pooled risk ratio of 1.17 with an interval of 0.40 to 3.40, with no heterogeneity. The post hoc analyses by pharmacological class put serotonin and noradrenaline reuptake inhibitors at the highest risk ratio, 2.93 with an interval of 0.97 to 8.82, whose lower bound comes close to 1 without crossing it. The authors point out that these groupings were not planned in the protocol.

Venlafaxine carries the only point estimate that stands out, a risk ratio of 4.53 with an interval running from 0.47 to 43.25. On its own this figure demonstrates nothing, and the width of the interval mainly measures how rare the events are. It does, however, have two features the authors point out: it goes in the direction expected from noradrenergic pharmacology, and venlafaxine is the only compound for which individual trials give consistent signals. It was studied as add-on therapy in two trials, one against paroxetine, the other alongside bupropion and sertraline, and in the authors’ words “Both trials reported a significantly increased risk of mania in participants treated with venlafaxine compared to other antidepressants.” This consistency does not amount to a demonstration at network level, it amounts to converging evidence: it makes a cautious position defensible without making it mandatory.

The P-score ranking puts amineptine first, then amitriptyline, then placebo, with venlafaxine in last position behind imipramine and desipramine. The spontaneous reading of this list has to be resisted. A P-score is a computed rank, not a measure of safety. The two best-ranked drugs are precisely those with the thinnest data, with confidence intervals among the widest in the network. This ranking therefore does not constitute a prescribing recommendation, and should not be cited as one.

ResultReading
Each antidepressant compared with placeboInconclusive
PEB readingNo risk ratio significantly higher than placebo, with estimates ranging from 0.31 to 4.53. Given the CINeMA confidence and the small number of trials comparing an antidepressant directly with placebo, this is an absence of demonstration, not a demonstration of absence.
Head-to-head comparisonsInconclusive
PEB readingNo antidepressant shows a switch rate significantly higher than another. The network cannot separate the drugs from one another.
VenlafaxineIsolated signal
PEB readingRisk ratio of 4.53, interval 0.47 to 43.25, and therefore compatible with no difference. Taken alone, the signal proves nothing, but this is the only compound for which the individual trials agree.
P-score rankingRelative rank
PEB readingAmineptine and amitriptyline at the top, placebo third, venlafaxine last. A statistical order internal to the network, with no value as a safety hierarchy and no value as a recommendation.
Switch beyond 10 weeksNo data
PEB readingNo included trial follows patients beyond this duration. The study can neither confirm nor rule out a delayed risk, a question the authors explicitly leave to future work.
RankTreatmentP-scoreRisk ratio versus placebo
1Amineptine0.83860.33 [0.04; 2.73]
2Amitriptyline0.79090.31 [0.01; 7.82]
3Placebo0.6364reference
4Bupropion0.57691.15 [0.08; 17.09]
5Paroxetine0.57421.13 [0.54; 2.37]
6Agomelatine0.55571.17 [0.40; 3.40]
7Fluoxetine0.52391.27 [0.43; 3.75]
8Tranylcypromine0.45251.66 [0.27; 10.25]
9Sertraline0.41102.03 [0.16; 25.04]
10Moclobemide0.36672.12 [0.42; 10.72]
11Desipramine0.35852.64 [0.08; 84.25]
12Imipramine0.24982.76 [0.80; 9.58]
13Venlafaxine0.16494.53 [0.47; 43.25]

None of these estimates excludes no effect. The width of the intervals, up to an upper bound of 84.25 for desipramine, reflects the number of events on which each comparison rests. Rank and risk ratio tell the same story: this ranking orders point estimates, it does not measure a difference in safety.

Critical appraisal

DomainJudgement
Power and confidence in the networkMajor weakness
PEB readingLow CINeMA confidence for 77 comparisons out of 78 and very low for the last one, only six trials with a direct comparison with placebo, intervals reaching 84.25. This is the limitation that governs all the others: the network does not have the information needed to separate the drugs.
Definition of the outcomeReservation
PEB readingYoung Mania Rating Scale thresholds that differ from one trial to the next, categorical criteria elsewhere, clinical judgement alone in five trials. Heterogeneous definitions pull the estimates towards no difference. The corresponding sensitivity analysis did not change the results.
Length of exposureNo data
PEB reading6.9 weeks on average, 10 weeks at most. The authors note that this period exceeds the two-week threshold set by the ISBD task force for a switch to count as treatment-emergent, but that it says nothing about delayed risk. What the study does not measure should not be reported as a reassuring result.
Representativeness of the drugsReservation
PEB readingThe network is dominated by older drugs and by trials several of which date from 1989 to 2002. Emerging treatments were excluded because no events were reported, and regimens starting an antidepressant together with an antipsychotic were ruled out by design. Transfer to contemporary prescribing is not straightforward.
Methodological conductSound
PEB readingPreregistration on the Open Science Framework with amendments reported, PRISMA-NMA reporting, risk of bias assessed with RoB 2, multiple sensitivity analyses, confidence assessed with CINeMA and published as it stands. Transparency here is at the expected level.
Fit between claim and evidenceCalibrated
PEB readingThe authors do not present the ranking as a safety hierarchy, describe their post hoc analyses as exploratory and set out the limitations. The risk of overinterpretation lies downstream, with readers and in secondary coverage.
IndependenceTo be qualified
PEB readingNo funding for the study, which is the important point. Individual declarations of interest, however, are substantial: eight of the fifteen authors declare no conflict, the other seven declare honoraria or consulting roles with a large number of companies, described, where the declaration specifies it, as outside or unrelated to the submitted work. Nothing points to an effect on the results, but describing this as work free of competing interests would be inaccurate.

From the ranking to the prescription

Several elements have to be recalled here, with the clear statement that they lie outside the study: a network meta-analysis compares drugs, it does not rule on their national authorisation conditions. These points also call for checking against official sources before any decision, since statuses can change.

First point, and on its own it illustrates what the ranking is worth: the top-ranked drug in a ranking is not necessarily one that can be prescribed. Amineptine, which comes first here, has not been marketed in France since the late 1990s, its withdrawal being linked to a potential for abuse, information to be checked against official sources. Its rank rests on a single Italian trial from 1996 comparing 11 patients with 11 patients. The second, amitriptyline, is available in France as an antidepressant, but its rank rests on a single German trial, with 22 patients on amitriptyline. Turning a P-score rank into a course of action therefore amounts here to prescribing on the strength of two small trials, one per drug.

Second point, which belongs to guidelines and not to regulation: international guidelines place antidepressants as a second-line option in bipolar depression, as add-on to a mood stabiliser or a second-generation antipsychotic rather than as monotherapy. The study provides indirect support here: two thirds of the patients included received the antidepressant as add-on therapy, and the authors stress that the monotherapy subgroup rests on fewer trials and therefore offers less robust evidence. The cautious position on monotherapy is not derived from this work, it predates it, and this work does not contradict it.

Third point, on safety of use: agomelatine requires monitoring of liver function, the arrangements for which are set by the product information and should be checked in the version in force where you practise. Its rank in a ranking based on switch risk alone says nothing about this tolerability profile.

Level of evidence

Scientific72/100
Editorial87/100

PEB appraisal: confidence is high in the conduct of the work and in the honesty of its reporting. Preregistration, adherence to PRISMA-NMA, the risk of bias assessment and above all the publication, without softening, of an unfavourable CINeMA assessment place this article among the syntheses that can be read without fear of embellishment. Confidence is also high in the negative finding correctly understood: the available randomised trials do not establish that any given antidepressant carries a greater risk of switch than placebo over an exposure of a few weeks. It is low, and the authors say so, for any comparative claim between drugs. It is nil for the risk beyond 10 weeks, for lack of measurement in the included trials, and nil too for the idea that this work would justify relaxing vigilance after starting an antidepressant in a patient with bipolar disorder. An article that properly documents an uncertainty does not reduce it, it makes it visible.

The colleague test

What an experienced colleague would say if you put this study to them in two minutes, between two consultations.

“ So they found nothing, and they were honest enough to say why: there were not enough trials to find anything. I am not going to use the ranking to choose a drug, especially when the first one in the table is no longer sold in France. Venlafaxine bringing up the rear does not surprise me, and it does not change what I already do. What stays with me most is that ten weeks is not how long I keep watch over a patient I have just added an antidepressant for. ”

What this means in practice: the work confirms that current caution rests on converging evidence and not on a demonstration, which does not disqualify it. It provides no argument for switching drugs, and it provides a solid argument against relaxing monitoring on the strength of a non-significant result.

What you can do with this

  • Do not cite this work as evidence that antidepressants do not cause switches. The accurate wording, usable in a team meeting as well as with a patient, is that the available trials are too few to tell the drugs apart.
  • Do not turn the ranking into a course of action. The top two ranks each rest on a single small trial, and the drug ranked first is no longer marketed in France, a point to check before any team discussion.
  • Keep the caution about venlafaxine in bipolar disorder, knowing that this position rests on converging signals, including two individual trials, and not on a difference demonstrated at network level. Present it as such rather than as a certainty.
  • Keep monitoring for switch well beyond the first weeks, precisely because the study stops at 10 weeks. Agree with the patient and those close to them on the signs that should prompt them to get back in touch.
  • Use this article for teaching: it illustrates the difference between absence of evidence and evidence of absence, and the trap of a ranking produced from data that cannot support it.
  • The course of action is set out in the NICE decision tree for depression in adults.

Frequently asked questions

Does this network meta-analysis show that antidepressants are safe in bipolar depression?

No. It shows that the available randomised trials do not detect a difference in switch risk between antidepressants and placebo, over an exposure of a few weeks. With low CINeMA confidence for 77 comparisons out of 78 and very low for the last one, this result reflects a lack of information, not established safety.

Should the top-ranked antidepressant be preferred to limit the risk of manic switch?

No, for two reasons. The first is methodological: a P-score rank does not measure safety, it orders estimates whose intervals overlap, and the top two ranks here each come from a single small trial. The second is practical: the drug ranked first, amineptine, is no longer marketed in France, information to be checked against official sources.

Does a risk ratio of 4.53 for venlafaxine justify avoiding it in bipolar depression?

That figure is not enough on its own, since its interval runs from 0.47 to 43.25. It does, however, support converging evidence: venlafaxine is the only compound in the network for which individual trials give consistent signals, two add-on trials having reported a significantly increased risk of mania compared with other antidepressants. Setting venlafaxine aside in bipolar depression when an alternative exists remains a defensible position, to be presented as reasoned caution and not as a demonstrated contraindication.

Why is trial duration a particular problem for antidepressant-induced mania?

Because the observation window, from 1.7 to 10 weeks depending on the trial, covers only the acute phase. The authors stress that the long-term risk remains uncertain and call for longer trials or observational data. The absence of events in this window is not an absence of risk, it is an absence of measurement.

How do differing definitions of manic switch affect the results?

They push towards no difference. Depending on the trial, switch is defined by a Young Mania Rating Scale threshold, ranging from above 10 to 16 or more, by categorical DSM criteria, by an adverse effects scale or by the clinician’s judgement. These definitions do not designate the same event. Pooling them adds noise to the signal. The authors tested this point in a sensitivity analysis, with no change in the results, while acknowledging that subtle inconsistencies may have gone undetected.

What does this study say about antidepressant monotherapy in bipolar depression?

Little, and the authors say so. Two thirds of the patients received the antidepressant as add-on to a mood stabiliser, and the monotherapy subgroup rests on fewer trials, and so on a less robust evidence base. Sensitivity analyses by treatment regimen did not change the results, but they are not enough to validate monotherapy, which international guidelines continue to advise against.

Annotated bibliography

1. Oliva V, De Prisco M, La Spina E, Paolucci S, Fico G, Anmella G, Hidalgo-Mazzei D, Murru A, Pompili M, Fornaro M, Solmi M, Yildiz A, Leucht S, Vieta E, Radua J. Switch to mania after acute antidepressant treatment for bipolar depression: a systematic review and network meta-analysis of randomised controlled trials. eClinicalMedicine 2025; 87: 103413. DOI 10.1016/j.eclinm.2025.103413. PMID 40823496. Frequentist random-effects network meta-analysis of 13 randomised trials, 1362 patients, 12 antidepressants compared with placebo and with each other on the risk of manic or hypomanic switch in bipolar depression. Protocol preregistered on the Open Science Framework, PRISMA-NMA reporting, RoB 2, CINeMA. No comparison concludes to a risk significantly higher than placebo, with low confidence for 77 comparisons out of 78 and very low for the last one. No funding. Main interest: establishing the real state of the evidence on an everyday practice question. Main limitation: the network lacks the power to separate the drugs, and mean trial duration is 6.9 weeks.

2. Vieta E, Martinez-Aran A, Goikolea JM, et al. A randomized trial comparing paroxetine and venlafaxine in the treatment of bipolar depressed patients taking mood stabilizers. J Clin Psychiatry 2002; 63(6): 508-512. DOI 10.4088/jcp.v63n0607. One of the two individual trials on which the venlafaxine signal rests, cited as such by the publication analysed. Thirty patients per arm, as add-on to a mood stabiliser.

3. Post RM, Altshuler L, Leverich G, et al. Mood switch in bipolar depression: comparison of adjunctive venlafaxine, bupropion and sertraline. Br J Psychiatry 2006; 189(2): 124-131. DOI 10.1192/bjp.bp.105.013045. The second individual trial behind the venlafaxine signal, 65 patients on venlafaxine, 51 on bupropion, 58 on sertraline, all as add-on therapy. It is the largest direct contributor to the network for this drug.

4. Yildiz A, Siafis S, Mavridis D, Vieta E, Leucht S. Comparative efficacy and tolerability of pharmacological interventions for acute bipolar depression in adults: a systematic review and network meta-analysis. Lancet Psychiatry 2023; 10(9): 693-705. DOI 10.1016/S2215-0366(23)00199-2. The reference network meta-analysis on the acute treatment of bipolar depression, from which the work analysed sets itself apart by restricting the network to antidepressants alone. It found more switches on antidepressants than with other classes, with no difference from placebo.

5. Pacchiarotti I, Bond DJ, Baldessarini RJ, et al. The International Society for Bipolar Disorders (ISBD) task force report on antidepressant use in bipolar disorders. Am J Psychiatry 2013; 170(11): 1249-1262. DOI 10.1176/appi.ajp.2013.13020185. The guideline framework against which the authors set their conclusions, notably on add-on use rather than monotherapy and on the higher risk in bipolar I disorder.

References 2 to 5 are cited from the reference list of the publication analysed. Their PubMed identifiers are not reproduced here.

Editorial collections

Topics

Verified on 12 August 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on 16 September 2026, against the figures of the French version and against the source. How we verify what we publish.
Content published by Psychiatry Evidence Base is produced according to the principles of evidence-based medicine. Every analysis rests on an independent critical reading of the scientific literature and aims to help health professionals interpret it. The information presented replaces neither official guidelines, nor clinical reasoning, nor individualised care. Medicine evolves continuously, and some data may change as new scientific evidence appears.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigour.

Report an error in this analysis

Follow Dr Stroescu on LinkedIn, for the review every Saturday