Published on 17 September 2026
Auditory hallucinations: what does a targeted psychotherapy actually change?
JAMA Psychiatry · 2026; advance online publication of 17 June 2026, no volume or pagination assigned as of the date of consultation · Fässler et al.
DOI 10.1001/jamapsychiatry.2026.1443
PMID 42307965
Scientific 73
Editorial 71
The essentials
Twenty-three randomised trials published between 2004 and 2025, 2,016 adult participants with a schizophrenia spectrum disorder, 1,081 in intervention arms spread across five therapeutic approaches and 935 in control arms; twenty-two trials provide usable data. At end of treatment, psychological and psychosocial interventions that specifically target auditory hallucinations reduce their severity compared with controls: standardised mean difference, Hedges’ g, −0.18 (95% CI: −0.31 to −0.05; p = 0.007; number needed to treat 16; GRADE, low-certainty evidence), with low heterogeneity (I² 31.8%; p = 0.06). The effect is maintained at follow-up, on average 27 weeks after inclusion: −0.15 (95% CI: −0.27 to −0.03; p = 0.01). In subgroup analysis, avatar therapy is the only one of the five approaches to cross the significance threshold on severity (−0.37; 95% CI: −0.51 to −0.22; p < 0.001). Secondary endpoints point the same way for voice-related frequency and distress, total symptoms, positive symptoms, delusions, depression and anxiety, but two endpoints do not move: negative symptoms (0.01; 95% CI: −0.13 to 0.14; p = 0.90) and malevolence beliefs attributed to the voices (−0.09; 95% CI: −0.24 to 0.07; p = 0.26). Dropouts do not differ between arms, which argues for good acceptability, but the collection of adverse events is poor: only six trials document them in detail, and eight report none at all. The authors conclude to efficacy with effect sizes ranging from small to medium, and to substantial differences between approaches. Three caveats govern how this should be read. The hierarchy between approaches comes from subgroup analyses, not from randomised head-to-head comparisons, and the authors say so themselves. Eleven of the twenty-two evaluated trials are at high risk of bias. Finally, against intuition, it is the low-risk-of-bias trials and the active-comparator trials that produce the largest effects, which rules out explaining this result by a simple lack of methodological rigour.
Context
The authors set out the clinical problem at the outset: auditory hallucinations are a cardinal feature of schizophrenia spectrum disorders, with a lifetime prevalence they place between 64% and 80%, and they are often accompanied by severe distress. Some of the people affected continue to hear voices despite antipsychotic treatment conducted according to the rules, which explains why an entire literature has grown up around the phrase medication-resistant hallucinations. The question posed here is therefore not to replace pharmacotherapy, but to find out what a psychotherapy built to act on the voice itself actually contributes.
The authors claim priority, stated in the discussion: to their knowledge, this is the first systematic review with meta-analysis covering the full range of psychological and psychosocial interventions specifically designed for auditory hallucinations. The claim is defensible for this broad scope, but it does not excuse situating the work. A 2020 Cochrane review had already examined avatar therapy and had not found a consistent effect, across three studies and 195 participants, with certainty levels ranging from very low to moderate depending on the endpoint; it does not appear among the article’s 72 references. A network meta-analysis published in 2026 in Psychological Medicine, which includes one of the senior authors of the work analysed here, specifically pitted avatar therapy against cognitive behavioural therapy; published after the manuscript’s acceptance, dated 4 April 2026, it could not be discussed in it. These three pieces of work do not tell the same story, and that is precisely what makes the exercise worthwhile.
The study at a glance
| Question (PICO) | |
|---|---|
| Population | |
| Adults with a schizophrenia spectrum disorder and current auditory hallucinations, with at least 75% of patients carrying this diagnosis in each arm. 23 randomised trials published between 2004 and 2025, 2,016 participants, of whom 1,081 in intervention arms and 935 in control arms; 22 trials provide usable data. Mean age 38.6 years (trial means from 25.9 to 46.0), 53.3% men, mean sample size per trial 87.6 (20 to 345) | |
| Intervention | |
| Any psychological or psychosocial intervention primarily aiming to reduce a feature of auditory hallucinations, with no restriction on care setting. The authors group them into five categories: cognitive behavioural therapy (5 trials), avatar therapy (6 trials, several of them in virtual reality), dialogue and relating-to-voices therapy (2 trials), self-help programmes, online or computerised, referred to as digital (5 trials), and integrative approaches combining several models (5 trials). Median trial duration 12 weeks (6 to 39) | |
| Comparator | |
| Active controls, treatment as usual, waiting list, or a combination of these. Actual distribution across the 23 trials: treatment as usual alone in 14 trials, enhanced supportive counselling in 3, a befriending intervention in 1, cognitive behavioural therapy in 2, waiting list combined with treatment as usual in 3 | |
| Endpoints | |
| Primary endpoint: between-group change in the severity of auditory hallucinations from baseline to end of treatment. Secondary endpoints: severity at follow-up, other features of the voices, other clinical endpoints at end of treatment and at follow-up, measured with validated scales, foremost among them the PSYRATS-AH. Acceptability measured by dropouts, adverse events as reported in each trial. Additional analyses of number needed to treat and grading of the certainty of evidence using the GRADE approach | |
| Design | |
| Systematic review with random-effects pairwise meta-analysis, Hedges’ g, subgroup analyses and mixed-effects meta-regression. Search of Embase, MEDLINE, PsycArticles, PsycInfo, PSYNDEX and the Cochrane Library, initial search on 18 November 2024, updated on 18 August 2025. 7,408 records identified, 109 full texts screened, 23 trials included. PRISMA-compliant. Protocol registered on PROSPERO, CRD42023475704, and published in PLOS One in July 2024 · CEBM 1a |
Quality control
A methodological note before reading this table. The article’s full text was consulted, including its three tables and the trial-by-trial description. The items below come directly from it. A single access caveat remains: the eighteen figures and eight electronic tables of the supplementary material, which carry in particular the detail of the GRADE grids, the Egger’s tests and the number-needed-to-treat calculation, were not opened. Where an item appears only in that supplement, this is stated.
| Criterion | Status |
|---|---|
| Prior registration | Solid |
| Finding PROSPERO CRD42023475704, with a protocol published in a peer-reviewed journal in July 2024, more than a year before the literature search closed. Deviations from the preregistration are declared and referred to an appendix of the supplement | |
| Literature search | Solid |
| Finding Six databases searched, including PsycInfo and PSYNDEX, initial search in November 2024 then an update on 18 August 2025. 7,408 records identified, 109 full texts screened, 23 trials retained. Selection, extraction and risk-of-bias rating carried out in duplicate, with a third assessor used when disagreement persisted. The scope covers German-language literature, often absent from Anglo-Saxon reviews | |
| Blinding of participants and therapists | Impossible by design |
| Finding No psychotherapy trial can conceal from the patient or the therapist what is being administered. This is not a shortcoming of the authors, it is a constraint of the field, and it holds for the whole body of evidence | |
| Assessor blinding | Documented trial by trial |
| Finding The article’s descriptive table gives the blinding status for each of the 23 trials: 14 are single-blind, that is, with masked assessment, 1 is partially masked, 5 are explicitly open, 2 do not report this point, and 1 relies only on self-assessments. Assessor blinding is therefore the rule rather than the exception, which substantially weakens the hypothesis of an effect driven by informed assessors | |
| Homogeneity of comparators | Caveat, but not the one expected |
| Finding Treatment as usual alone dominates, 14 trials out of 23, against 6 active-comparator trials and 3 waiting-list-plus-treatment-as-usual trials. Superiority over treatment as usual and superiority over an active intervention are not the same demonstration. The type of control does moderate the result (Q = 10.23; p = 0.02), but in the opposite direction from intuition: only the active-comparator trials reach significance | |
| Nature of the comparisons | Caveat |
| Finding This is a random-effects pairwise meta-analysis, with each approach compared against its own controls, and the approaches then compared with each other by mixed-effects meta-regression. It is not a network meta-analysis: no formal indirect comparison is estimated, and no result comes from a network model | |
| Measurement instruments | Caveat |
| Finding The article names the instruments trial by trial. The PSYRATS-AH, a clinician-rated interview, serves as the severity measure in almost all the trials, most often supplemented by the BAVQ-R, a patient-completed beliefs questionnaire. These two formats do not respond in the same way to unblinding. The secondary endpoints rest on more heterogeneous scales, which the authors count among their limitations | |
| Grading of certainty | Caveat |
| Finding The body of the article gives two explicit levels, low for the primary endpoint and also low for the avatar subgroup. The complete GRADE grids are referred to an electronic table in the supplement, not consulted. No level is therefore readable, in the body of the article, for each of the secondary endpoints taken individually | |
| Risk of bias | Weak point of the evidence base |
| Finding Of the 22 trials evaluated with the RoB 2 tool, 6 are at low risk, 5 raise some concerns and 11 are at high risk. The authors make this one of their main limitations. Risk of bias significantly moderates the result (Q = 9.59; p = 0.02) | |
| Heterogeneity | Low |
| Finding I² 31.8% on the primary endpoint (p = 0.06), and a jackknife analysis finds no single study responsible for the heterogeneity. It remains more pronounced for voice frequency (I² 47.4%; p = 0.006) and null on several secondary endpoints | |
| Publication bias | Sought, not found |
| Finding Visual inspection of funnel plots and Egger’s test on the primary endpoint: z = −0.94; p = 0.35. No signal on the secondary endpoints either. Nothing therefore argues for publication bias. This is not a blank cheque: across 19 observations, Egger’s test has limited power and a non-significant result does not rule out a small-study effect | |
| Independence | Declared, and worth knowing |
| Finding Funding: an Elsa Neumann doctoral scholarship from the state of Berlin, with the funders declared to have played no role. Two disclosures deserve the reader’s attention. Stefan Leucht declares personal fees from twelve pharmaceutical companies, unrelated to the submitted work. Kerem Böge, the senior author, declares being a co-founder of three health companies, Kiso GmbH, Mental Hub and FIL, and having received consulting or speaking fees from Boehringer Ingelheim, Clicks Therapeutics and Angelini. In a synthesis whose most visible result favours a technology-based intervention, this position should be known. In addition, one of the included trials has as its first author the review’s own first author, and the authors specify that another assessor rated that trial’s risk of bias in her place. Finally, the authors disclose the use of a language model between July 2025 and February 2026 for language checking, R code review and shortening sections | |
| Tolerability and acceptability | Acceptability yes, tolerability poorly documented |
| Finding Study dropout proportions are comparable, 15.0% in the intervention arms and 16.2% in the control arms, with no difference in risk (relative risk 0.91; 95% CI: 0.75 to 1.11; p = 0.36). By contrast, only six trials assess and report adverse events in detail and eight report none at all. The authors say so themselves: the collection of adverse effects is insufficient and centred on serious events, at the expense of milder but clinically relevant ones | |
Results
| Endpoint | Published result |
|---|---|
| Severity of auditory hallucinations, end of treatment | SMD −0.18 (95% CI: −0.31 to −0.05); p = 0.007; I² 31.8%; NNT 16; GRADE, low certainty |
| PEB readingStatistically present, modest in size the effect is real in the statistical sense, and its size falls below the threshold usually used to describe a small effect | |
| Severity of auditory hallucinations, at follow-up | SMD −0.15 (95% CI: −0.27 to −0.03); p = 0.01; I² 5.3%. Mean follow-up 27.2 weeks, from 12 to 78 weeks depending on the trial |
| PEB readingThe effect does not fade the size barely decreases and the interval still excludes zero. This is an important point for a psychotherapy, even though few trials extend beyond twelve months | |
| Avatar therapy subgroup | 6 trials, 7 observations. Severity at end of treatment −0.37 (95% CI: −0.51 to −0.22); p < 0.001; I² 0.01%. At follow-up −0.26 (95% CI: −0.41 to −0.11); p < 0.001. Frequency −0.34, distress −0.29, both significant. Malevolence beliefs −0.09 (95% CI: −0.32 to 0.14), not significant. GRADE, low certainty |
| PEB readingThe largest effect, but in a subgroup the six trials are Craig 2018, Percie du Sert 2018, Dellazizzo 2021, Liang 2022, Garety 2024, which contributes two observations, and Smith 2025. Four of them use an active comparator and three have masked assessment. The authors acknowledge that this subgroup concentrates the active comparators and the low-risk-of-bias trials, which makes its apparent superiority inseparable from its methodological characteristics | |
| The other four approaches | Severity at end of treatment: cognitive behavioural therapy 0.05 (95% CI: −0.18 to 0.28); dialogue and relating-to-voices therapy 0.14 (95% CI: −0.40 to 0.68); digital approaches −0.05 (95% CI: −0.37 to 0.26); integrative approaches −0.25 (95% CI: −0.82 to 0.31). None is significant |
| PEB readingNot significant is not zero each subgroup rests on 2 to 4 trials, substantially underpowered, with very wide intervals. The authors write it themselves: these results may reflect a lack of power, not an absence of effect. They also note that cognitive behavioural therapy has proven itself on positive symptoms in general, which is not the question asked here | |
| Secondary endpoints on the voices | Frequency −0.28 (95% CI: −0.45 to −0.10); p = 0.002. Distress −0.19 (95% CI: −0.34 to −0.04); p = 0.01. Malevolence beliefs −0.09 (95% CI: −0.24 to 0.07); p = 0.26 |
| PEB readingTwo gains, one miss frequency moves more than overall severity, distress follows, but what the patient believes about the malevolence of the voice does not change. This is an interesting finding in itself: these therapies appear to act on the presence and lived experience of the voice more than on the belief system surrounding it | |
| Clinical secondary endpoints | Total symptoms −0.29 (95% CI: −0.48 to −0.11); positive symptoms −0.24 (95% CI: −0.38 to −0.11); delusions −0.27 (95% CI: −0.41 to −0.12); depression −0.19 (95% CI: −0.30 to −0.07); anxiety −0.28 (95% CI: −0.48 to −0.08). Negative symptoms 0.01 (95% CI: −0.13 to 0.14); p = 0.90 |
| PEB readingConsistency of direction, with one clear exception five of six endpoints move in the same direction, with magnitudes close to that of the primary endpoint or slightly larger. Negative symptoms do not move at all, which is consistent with the idea of an intervention targeted at a symptom rather than at the disorder | |
| Moderators and sensitivity analyses | The result is moderated by type of control (Q = 10.23; p = 0.02), risk of bias (Q = 9.59; p = 0.02) and therapeutic approach (Q = 26.07; p < 0.001), but not by treatment duration (Q = 2.55; p = 0.11). Restricted to low-risk-of-bias trials, the effect becomes −0.27 (95% CI: −0.44 to −0.09). Restricted to active-comparator trials, it becomes −0.32 (95% CI: −0.50 to −0.15). Excluding the waiting-list trials and the Chinese trial does not change the conclusions |
| PEB readingThe study’s most counter-intuitive result one would expect methodological rigour to shave the effect down; instead it increases it. The explanation the authors propose is structural: active comparators occur almost exclusively in the avatar trials, which are themselves mostly at low risk of bias. These three moderators are therefore confounded with one another and cannot be disentangled in this dataset | |
| Publication bias | Egger’s test on the primary endpoint: z = −0.94; p = 0.35. No signal on the secondary endpoints |
| PEB readingSought, not found the field’s classic argument, of small trials run by the teams that designed the intervention, is not borne out here. The test’s power remains low across 19 observations | |
| Tolerability and acceptability | Study dropout 15.0% in intervention versus 16.2% in control (relative risk 0.91; 95% CI: 0.75 to 1.11; p = 0.36). Treatment dropout 17.3% versus 12.2%, with no difference in the paired comparison (relative risk 0.97; 95% CI: 0.66 to 1.43; p = 0.87). Six trials report adverse events in detail, eight report none |
| PEB readingSolid acceptability, poorly documented tolerability patients stay, which is in itself a favourable signal and the strongest argument for offering these approaches. But a body of evidence where eight trials out of twenty-three report no adverse event at all does not allow one to claim good tolerability: it only allows one to say that no danger signal was reported | |
Two notions deserve to be spelled out, because they determine the scope of the result. The standardised mean difference is not a number of points on a scale: it is a gap between two means expressed relative to the dispersion of the measurements, here in the form of Hedges’ g, which allows different instruments to be pooled but forbids translating the figure directly into a clinical gain. By the usual interpretive conventions, 0.2 marks the threshold of a small effect; the pooled estimate remains below it. The avatar subgroup, at 0.37, sits between a small and a medium effect, which matches the wording the authors chose, effect sizes ranging from small to medium.
The number needed to treat of 16 calls for the same caution. Such an index is calculated from a proportion of responders. Here, the primary endpoint is continuous, which requires defining a response threshold and using an assumption about the response rate in the control group. The body of the article details neither this threshold nor this assumption, and refers the calculation to an electronic table in the supplement, not consulted. The figure remains useful for gauging the order of magnitude; it should not be handled as though it came from a direct count of improved patients.
Finally, a comparison that is not in the article but that sheds light on its most visible result. The network meta-analysis published in 2026 in Psychological Medicine, across 26 trials and 2,273 participants, directly pitted avatar therapy against cognitive behavioural therapy in patients with medication-resistant hallucinations. At end of treatment, it finds no difference in hallucination severity, with a standardised mean difference of −0.23 whose 95% confidence interval runs from −0.55 to 0.10, that is, it crosses zero. Avatar’s advantage there appears at three months, −0.37 (95% CI: −0.69 to −0.05), and on overall psychotic symptoms, −0.41 (95% CI: −0.75 to −0.06). The authors of this second work themselves report a small-study effect, with a significant Egger’s test below the 0.01 threshold for hallucination severity, and an overall risk of bias ranging from moderate to high. Two lessons follow. First, avatar’s superiority over cognitive behavioural therapy is not established in a head-to-head comparison at end of treatment. Second, the two pieces of work converge on durability: avatar’s advantage shows up at a distance from treatment, at three months in the network meta-analysis, at a mean follow-up of 27 weeks in the pairwise meta-analysis. Convergence is not confirmation; the two bodies of evidence overlap substantially.
Critical appraisal
| Domain | Judgement |
|---|---|
| Level of evidence of the synthesis | Solid |
| Finding Top of the hierarchy: randomised trials only, a preregistered and then published protocol, six bibliographic databases, duplicate selection and extraction, PRISMA compliance, GRADE grading, a publication-bias test. The apparatus is what one expects of a major review | |
| Who assesses, and under what blind | Caveat largely resolved |
| Finding In a psychotherapy trial, patient blinding does not exist; only assessment can be masked, and this is therefore where robustness is judged. The article documents this point trial by trial: 14 of the 23 trials have masked assessment, 5 are open, 1 is partially masked, 2 are unreported, 1 relies on self-assessment alone. Above all, the analysis restricted to low-risk-of-bias trials, including the outcome-measurement domain, does not reduce the effect but increases it, from −0.18 to −0.27. The hypothesis of an estimate inflated by informed assessors therefore does not hold up against this data | |
| Patient expectation and researcher allegiance | Caveat maintained |
| Finding A therapy that stages the patient’s voice generates strong expectations, and research-team allegiance is known to inflate effects in psychotherapy trials. The article does not analyse this factor. The signal nonetheless exists in the literature: the 2020 Cochrane review noted that, in one of the avatar trials, authors had taken part in developing the system being evaluated, and that in another, patents were being filed. This argument is weakened here, though not cancelled out, by the fact that the active-comparator trials give the largest effects | |
| Mixed comparators | Caveat, in the opposite direction |
| Finding Pooling superiority over waiting list with superiority over an active intervention amounts to mixing two distinct clinical questions: is it better than nothing, and is it better than something else. The article separates the two, and the result is surprising: only the active-comparator trials reach significance (−0.32; 95% CI: −0.50 to −0.15), against nothing significant for treatment as usual alone or the waiting list. This is not a demonstration of specificity, since the authors themselves note that the active comparators are concentrated in the avatar trials. Type of control, approach and risk of bias are confounded | |
| Subgroup, not a randomised comparison | Main caveat |
| Finding The hierarchy between approaches rests on subgroup analyses within the same pairwise meta-analysis. This type of analysis does not preserve randomisation between approaches: the avatar trials differ from the cognitive behavioural therapy trials in their populations, their controls, their dates and their methodological quality. Nor is it a network meta-analysis, so no formal indirect comparison is offered. The authors explicitly list this among their limitations: subgroup superiority must be interpreted with caution in the absence of head-to-head trials | |
| Underpowered subgroups | Caveat |
| Finding The five subgroups rest on 2 to 6 trials. That cognitive behavioural therapy comes out at 0.05 across three trials does not demonstrate its inefficacy on the voices, it demonstrates that this is unknown. Reading this table as a ranking of efficacy would be a reasoning error, the one that treats an absence of significance as evidence of absence | |
| Result diverging from prior evidence | Caveat |
| Finding The 2020 Cochrane review found no consistent effect of avatar therapy, across three studies and 195 participants, all at high risk of bias on at least one domain, with short-term data only. Six years later, a subgroup of six trials gives the largest effect observed. The accumulation of new trials, including the 2024 AVATAR2 trial and the 2025 Danish trial, probably explains this shift, but a reversal of conclusion calls, as a matter of principle, for verification, not endorsement | |
| Small-study effect | Sought, not found here |
| Finding The field is made up of small trials run by the teams that designed the interventions, and the 2026 network meta-analysis objectively finds a small-study effect on this same endpoint (Egger’s test, p < 0.01). The work analysed here applies the same test and does not find it (z = −0.94; p = 0.35). Two closely related bodies of evidence, two opposite results on this point: the difference in scope, a broader body of evidence here, restricted there to the avatar-versus-cognitive-behavioural-therapy comparison, probably explains the discrepancy, but the doubt is not resolved | |
| External validity | Caveat |
| Finding The trials come almost entirely from the United Kingdom, Northern Europe, Canada and Australia, with a single Chinese trial; no French trial. The interventions tested presuppose a trained therapist, a structured protocol and, for avatar, technical equipment and specific support. None of this can be inferred from an effect size. The authors further note that concomitant pharmacological treatment is rarely described, which makes it impossible to know what medication background these effects are added on top of | |
| Tolerability | Acceptability yes, tolerability not demonstrated |
| Finding Dropouts are identical in both arms, which argues for good acceptability and allows, in practice, offering these approaches despite a modest effect size. But the authors themselves describe the collection of adverse events as limited and insufficient, centred on serious events. Good measured acceptability is not the same as demonstrated good tolerability | |
Let us separate what this data allows one to say. What is demonstrated, with a certainty the authors themselves describe as low: across 23 randomised trials, psychotherapies built to target the voice do better than control conditions on hallucination severity at end of treatment, with a modest magnitude; the effect is maintained at follow-up; it extends to voice frequency, positive symptoms, delusions, depression and anxiety; and acceptability is good. What is suggested: avatar therapy stands out from the other approaches, and the robustness of the result across sensitivity analyses argues for a real effect rather than a measurement artefact. What is not demonstrated, and what a hurried reader risks wrongly concluding: that the other approaches are ineffective, when their subgroups are simply too small; that avatar is superior to cognitive behavioural therapy, which would require a head-to-head comparison; and that these interventions are well tolerated, which would require that adverse events had actually been sought, something eight trials out of twenty-three did not do. What amounts to expert opinion, including our own: the fact that the most rigorous trials and the most demanding comparisons give the largest effects is this work’s most intriguing result. It may signal a genuinely specific intervention, or merely the fact that the best studies in the evidence base are also the ones testing the most recent and best-funded approach. A trial directly comparing avatar therapy with a voice-targeted cognitive behavioural therapy, with masked assessment, would settle the matter.
Level of evidence
PEB assessment: low confidence, in the sense of the grading the authors themselves adopt, in the existence of a modest benefit of targeted psychotherapies on the severity of auditory hallucinations, at end of treatment and at follow-up; also low confidence, in the same sense, in the effect of avatar therapy taken on its own, but very low confidence in its relative superiority, which rests on a subgroup analysis confounded with comparator type and risk of bias, and which is not found in a head-to-head comparison at end of treatment in an independent piece of work; reasonable confidence in acceptability, but insufficient evidence on tolerability. The review’s method is of good quality and the authors’ conclusion is well calibrated, with an explicit mention of the low certainty in the abstract itself, which is not so common. What limits the scope lies in the material: eleven of twenty-two trials at high risk of bias, subgroups of two to six trials, patchy collection of adverse events, and a structure in which the three variables that explain the result, therapeutic approach, comparator type and methodological quality, vary together. The authors identify these limitations themselves, which speaks in their favour.
The colleague test
What an experienced colleague would say if you summarised this work to them between two consultations.
“ So the voice-hearing psychotherapies work a little, the effect is small, it holds up over time, and the certainty stays low. Avatar stands out from the pack, but in a subgroup, not in a trial that actually pits it against the others. And when you look closely, the more rigorous studies find more effect, not less: that part is rather reassuring. What I mainly take from it is that the patients stay. Still, eight trials out of twenty-three didn’t even look for adverse events. ”
What this means in practice: for any psychotherapy trial, two questions are worth more than one. Compared with what, first, because here the comparator changes everything. Then, who assessed it, and did they know, because that is the only blind available in this field. And a third, almost always forgotten: was anyone looking for what could go wrong. A psychotherapy with no adverse events reported is not a psychotherapy without adverse events.
What you can take from this on Monday morning. This work does not create a new prescription, but it provides language, an order of magnitude and a course of action.
- Treat the voice as a target in its own right. The literature this review rests on exists because patients continue to hear voices under well-conducted treatment. Explicitly naming this residual symptom in consultation, rather than immediately filing it under pharmacological failure, opens a useful conversation.
- Measure before concluding. The included trials overwhelmingly rely on two complementary formats: a clinician-rated interview, the hallucinations subscale of the PSYRATS, and a patient-completed beliefs questionnaire, the BAVQ-R. Documenting severity, frequency and distress at baseline makes it possible to know later whether anything happened, something no global impression can provide.
- State the right order of magnitude, and the right scope. One can honestly tell a patient that a targeted psychotherapy brings, on average and on low-certainty data, a real but modest benefit, which appears to hold for several months, and that people who start this care tend to continue it. One can add what these therapies do not do: they change neither negative symptoms nor beliefs about the malevolence of the voice. This is a reasonable proposal, not a promise.
- Do not oversell avatar. It is the most promising approach in this synthesis, it is not an approach validated against its competitors, and its apparent superiority is inseparable from the fact that its trials are also the best constructed in the evidence base. A search of the ClinicalTrials.gov registry on 11 August 2026 found no trial of avatar therapy proper conducted in France. The one related French registration identified on that registry concerns something else: a multicentre trial sponsored by the Guadeloupe university hospital centre, whose only declared site is in Les Abymes and whose inclusion criteria also mention recruitment at the Nîmes university hospital centre, evaluating a fifteen-week integrative group programme of which a single session is devoted to creating a voice avatar, not yet recruiting as of the date of consultation. A trial registry does not describe routine care on offer: this source does not document a French implementation of avatar therapy, and it does not prove that none exists. Readers outside France should expect a comparably patchy picture in their own registries and verify it directly rather than assume this example transfers.
- Keep referring patients to the psychosocial approaches available. The pooled result covers the full set of targeted interventions, not avatar therapy alone, even though the avatar trials carry a substantial share of it. None of the other four approaches is disqualified by this synthesis: they are simply less studied. Structured psychotherapeutic support, within a service or a network specialised in psychotic disorders, remains an option consistent with this data.
- Take away the reading grid. Standardised mean difference on one side, the conventional threshold on the other, the nature of the comparator third. And a fourth reflex, which this work illustrates well: when an author writes that one approach is superior, check whether the comparison is randomised or whether it rests on placing two subgroups side by side. These reflexes are useful for any psychotherapy trial, far beyond auditory hallucinations.
- The course of action is set out in the NICE decision tree for schizophrenia in adults.
Frequently asked questions
Should a targeted psychotherapy be offered to a patient who still hears voices on antipsychotic treatment?
This data supports the proposal without making it mandatory. The average benefit is small in magnitude and the certainty of evidence is judged low by the authors, but it is maintained at follow-up and acceptability is good, with dropouts identical in both arms. Tolerability, however, is poorly documented: eight trials out of twenty-three report no adverse event at all, which does not mean none occurred. In a situation where the alternative is often a dose increase or antipsychotic combination, the balance between expected benefit and risk incurred nonetheless tips toward offering it.
Is avatar therapy superior to cognitive behavioural therapy?
This work does not demonstrate it. It shows that the six-trial avatar subgroup is the only one to reach significance against its own controls, while the three-trial cognitive behavioural therapy subgroup does not (0.05; 95% CI: −0.18 to 0.28). Comparing these two figures is not a randomised comparison: the trials differ in their comparators, their quality and their power, and three trials are not enough to conclude an absence of effect. The authors themselves list this caveat among their limitations. The network meta-analysis published in 2026 in Psychological Medicine, which addresses this question directly, finds no significant difference in hallucination severity at end of treatment, and an advantage at three months whose confidence interval skims zero.
Why was the 2020 Cochrane review negative on avatar?
It had three studies and 195 participants available, all at high risk of bias on at least one domain, with certainty levels mostly very low to low and short-term data only. The body of evidence has grown since, notably with the 2024 British AVATAR2 trial and the 2025 Danish trial, which carry substantial weight in the subgroup. A change of conclusion driven by the accumulation of data is how evidence-based medicine normally works, provided the new data is of better quality. On this specific point, it is: the avatar trials are mostly at low risk of bias and use an active comparator, which was not the case in 2020.
What does a standardised mean difference of −0.18 actually mean for a patient?
No direct translation is possible. This index expresses a gap between means relative to the dispersion of the measurements, which allows different scales to be pooled but forbids converting it into scale points or fewer hours of voice-hearing. By the usual conventions, the value remains below the threshold for a small effect. The number needed to treat of 16 helps convey the order of magnitude, provided one knows it derives from a continuous endpoint and therefore rests on a response threshold whose detail is referred to the supplementary material.
Does the fact that the patient knows what they are receiving invalidate the study?
No, and the article provides the means to check this. In a psychotherapy trial, blinding the participant is impossible; what remains possible, and becomes decisive, is blinding the assessor. Fourteen of the twenty-three included trials have masked assessment, five are open, one is partially masked, two do not specify, and one relies only on self-assessments. Above all, the analysis restricted to low-risk-of-bias trials, which includes the outcome-measurement domain, gives a larger effect, not a smaller one. The intuitive idea that the measured effect might be inflated by informed assessors is therefore not supported by this data.
Are these interventions available anywhere today?
The sources consulted do not allow a description of the French care offer specifically, and none of the twenty-three included trials was conducted in France. The ClinicalTrials.gov registry, searched on 11 August 2026, lists no trial of avatar therapy proper conducted in France, and only one related French registration: a trial sponsored by the Guadeloupe university hospital centre, whose declared site is in Les Abymes and whose inclusion criteria provide for recruitment at the Nîmes university hospital centre, covering a group programme that includes one avatar-creation session out of fifteen, not yet recruiting. A trial registry does not report on routine care, nor on its funding or reimbursement. The picture is likely similarly patchy in most other jurisdictions, and readers should check their own national registries and care pathways rather than assume any particular level of access.
Annotated bibliography
Fässler L, Koop S, Opper F, Bighelli I, Leucht S, Sabé M, Bajbouj M, Knaevelsrud C, Böge K (2026). Psychological Interventions Targeting Auditory Hallucinations in Persons With Psychotic Disorders: A Systematic Review and Meta-Analysis. JAMA Psychiatry, advance online publication of 17 June 2026; no volume or pagination assigned as of the date of consultation. DOI 10.1001/jamapsychiatry.2026.1443 · PMID 42307965 · PMCID PMC13276665. Source study analysed here. Free access under a CC BY-NC-ND licence. Full text consulted, including the three tables, the trial-by-trial description and the author information; only the supplementary material, eighteen figures and eight electronic tables, was not opened. Manuscript accepted on 4 April 2026, published online on 17 June 2026, with no volume or pagination assigned as of the date of consultation, with no erratum or correction attached. Funding: an Elsa Neumann doctoral scholarship from the state of Berlin, with no role declared for the funders. Declared conflicts of interest: personal fees from twelve pharmaceutical companies for Stefan Leucht, unrelated to the submitted work; for Kerem Böge, the senior author, co-founder status at three health companies, Kiso GmbH, Mental Hub and FIL, and consulting or speaking fees from Boehringer Ingelheim, Clicks Therapeutics and Angelini. The authors also disclose having used a language model for language checking, R code review and the shortening of sections.
Fässler L, Bighelli I, Leucht S, Sabé M, Bajbouj M, Knaevelsrud C, Böge K (2024). Targeted psychological and psychosocial interventions for auditory hallucinations in persons with psychotic disorders: Protocol for a systematic review and meta-analysis. PLOS One, 19(7), e0306324. DOI 10.1371/journal.pone.0306324 · PMID 38959279. Preregistered protocol of the review analysed here, published more than a year before the search closed. Useful for gauging the gap between intention and execution: the final article declares deviations from the preregistration and refers them to an appendix in the supplementary material. Doctoral funding declared in the protocol: an Elsa Neumann scholarship from the state of Berlin and support from the Charité’s Medical Scientist Program, this second funder no longer appearing in the 2026 article. A protocol describes an intention, it does not guarantee what was actually done.
Hsu TW, Liang CS, Changchien TC, Tseng PT, Carvalho AF, Stubbs B, Thompson T, Böge K, Hsu CW, Yang FC, Tu YK, Lin YH (2026). AVATAR versus cognitive-behavioral therapy for medication-resistant auditory hallucination: a systematic review and network meta-analysis. Psychological Medicine, 56, e107. DOI 10.1017/S0033291726104127 · PMID 41969063. An indispensable counterpoint, and this time a network meta-analysis, so combining direct and indirect comparisons. Across 26 trials and 2,273 participants, no significant difference between avatar and cognitive behavioural therapy on hallucination severity at end of treatment, an advantage for avatar at three months and on overall psychotic symptoms, with an objectified small-study effect and a risk of bias ranging from moderate to high. Kerem Böge is an author on both pieces of work, which makes the divergence all the more instructive. Published after the JAMA Psychiatry manuscript’s acceptance, it is not discussed there.
Aali G, Kariotis T, Shokraneh F (2020). Avatar Therapy for people with schizophrenia or related disorders. Cochrane Database of Systematic Reviews, 2020(5), CD011898. DOI 10.1002/14651858.CD011898.pub2 · PMID 32413166. The prior state of the question: little to no consistent effect of avatar therapy, across three studies and 195 participants, all at high risk of bias on at least one domain, with certainty levels ranging from very low to moderate depending on the endpoint, and short-term data only. The authors also noted that, in one of the trials, authors had taken part in developing the system being evaluated, and that in another, patents were being filed. This review is not cited in the 2026 article. Worth reading before taking the 2026 result as settled.
ClinicalTrials.gov, NCT07266376. Validation of a VOICE MAnagement Program in Schizophrenia. clinicaltrials.gov. A randomised controlled trial described as multicentre, sponsored by the Guadeloupe university hospital centre, with the Sud-Ouest Outre-Mer interregional clinical research and innovation group as a partner. Only one site is declared in the registry entry, in Les Abymes, but the inclusion criteria also provide for recruitment at the Nîmes university hospital centre. A fifteen-week integrative group programme of 90-minute weekly sessions, the eleventh of which is devoted to creating a voice avatar, compared with a media workshop; 116 participants planned, status not yet recruiting as of the date of consultation. The only related French trace identified on this registry. A source on implementation, not a source on efficacy.
What was consulted. Full text of the source study consulted on 11 August 2026 in its open-access archive version, tables and author information included; all figures cited come directly from it. The supplementary material, eighteen figures and eight electronic tables, was not opened, and the rare items that depend on it are flagged as such in the body of the analysis. References verified on Europe PMC, Crossref and OpenAlex on 11 August 2026, with no erratum or correction attached. Prior state of the question verified on the Cochrane Library and Europe PMC on 11 August 2026. Implementation of the interventions verified on the ClinicalTrials.gov registry on 11 August 2026. This analysis underwent an independent double reading.
