Published on 20 September 2026
Before the gaokao: a writing intervention that missed its own primary outcome
The essentials
This cluster randomized trial assigned fourteen senior high school classes, 587 students in all, in the hundred days before China’s national university entrance exam, to three daily twenty-minute sessions of guided narrative writing or to three neutral writing sessions of the same length. On the primary outcome, test anxiety, there is no difference: in the full sample every condition-by-time interaction is non-significant, and the authors state this plainly in their conclusion. Secondary outcomes show small effects favoring guided writing, with interaction effect sizes of d = 0.14 for depression, 0.16 for general anxiety, 0.14 for insomnia, and 0.23 for the composite distress score. The abstract reports values of 0.35, 0.37, and 0.23 for the intervention phase that are in fact the between-group differences measured at the end of treatment, a different quantity. The document reviewed is the accepted manuscript, released by the journal ahead of final production.
Context
Psychological distress tied to high-stakes exams is widespread and rarely addressed. China’s gaokao has its counterparts elsewhere, from graduation exams to competitive university admissions. Brief school-based interventions are multiplying, and their effects are generally small.
What makes this trial worth reading is not its result, it is its configuration. A negative primary outcome, stated without evasion by its own authors, sitting behind a title that foregrounds three secondary outcomes. It is exactly the kind of study a hurried reader gets backwards, and exactly the kind worth learning to spot.
The study at a glance
| Element | Content |
|---|---|
| Population | Senior-year students preparing for the gaokao, ages 16 to 20 |
| DetailMean age 18.23, 56.4% female, boarding residency required for inclusion, a single public high school. Prior diagnosed mental disorder excluded: this is not a clinical population. 59.3% had elevated test anxiety at baseline, defined as a TAI score above 35. | |
| Intervention | Three daily twenty-minute sessions of guided writing |
| DetailDelivered on smartphone, preceded by a standardized video. Day 1, thoughts and feelings about the upcoming exam. Day 2, negative emotions, their sources and impact. Day 3, positive emotions and personal growth. n = 290. | |
| Comparator | Neutral writing, an active comparator |
| DetailA factual, chronological account of the day, with no emotional content. Same duration, same device, same schedule. n = 297. This choice controls for attention and time spent, and it rules out any comparison against no intervention at all. | |
| Outcomes | Test anxiety as the primary outcome |
| Detail20-item TAI scale, with its worry and emotionality subscales. Secondary: 25-item RCADS, fifteen items for general anxiety and ten for depression, and the ISI for insomnia. Six measurement points, with a 15-day follow-up. | |
| Design | Two-arm cluster randomized controlled trial |
| DetailRandomization at the class level by an independent researcher, computer-generated sequence, arms labeled neutrally, principal investigator and statisticians blinded until analysis. Fourteen of sixteen eligible classes were randomized, the other two withdrawn by their homeroom teacher before the draw. Registered as ChiCTR2200058881. Funded by Chinese public foundations that the authors state had no role in the study, with no declared conflicts of interest. Oxford level of evidence 1b for the design. | |
Results
| Result | Value |
|---|---|
| Primary outcome, test anxiety | Not significant, all p above 0.10 |
| ReadingTotal score and both subscales, during the intervention phase and at follow-up. The authors conclude that efficacy is not demonstrated on this outcome. | |
| Secondary outcomes, interaction effect sizes | Depression d = 0.14; general anxiety d = 0.16; insomnia d = 0.14; RCADS total d = 0.23 |
| ReadingThese are the values that correspond to the phrase “during the intervention.” They are small. Corresponding p values are 0.007, 0.001, 0.035, and 0.001. | |
| Between-group differences at end of treatment | Depression d = 0.35; anxiety d = 0.37; insomnia d = 0.23; RCADS d = 0.39 |
| ReadingA different quantity from the row above. The abstract attributes these values to the intervention phase, while the body of the text correctly ties them to the end of treatment. The ratio between the two quantities ranges from 1.6 for insomnia to 2.5 for depression. | |
| At 15 days | RCADS d = 0.10, p = 0.049; insomnia p = 0.087, not significant |
| ReadingThe gap on the composite score narrows sharply, from 0.39 to 0.10, and the gap on insomnia is no longer significant. The authors do report that the between-group difference is maintained at this time point for depression and general anxiety, without an accompanying figure, and that no condition-by-time interaction appears during the follow-up phase. They describe a short-term buffering effect rather than a durable benefit. | |
| High-anxiety subgroup | Exploratory, n = 348 |
| ReadingA significant interaction during the intervention on the worry subscale alone, d = 0.18. The exploratory status is stated by the authors in both the abstract and the conclusion, which is worth noting. | |
| Retention and power | 82.5% at end of treatment, 70.9% at 15 days |
| ReadingAttrition was somewhat higher in the guided-writing arm, three to four points at each time point. Power was calculated a priori at 85% for a d of 0.25, targeting 576 participants, with 587 enrolled. The document reviewed carries two different figures for the 15-day follow-up, 416 participants in the body of the text and 429 in the appendix. | |
Critical appraisal
| Domain | Judgment |
|---|---|
| Randomization | Low risk |
| FindingComputer-generated sequence, independent operator, neutrally labeled arms, statisticians blinded until final analysis. Recruitment preceded randomization, so there is no recruitment bias within clusters. Rare, and well done. | |
| Outcome measurement | High risk |
| FindingAll four outcomes are entirely self-reported and participants were not blinded. The active comparator mitigates this problem without removing it. | |
| Accounting for clustering | Insufficient |
| FindingOnly fourteen clusters, and the primary model ignores this, the authors writing that a three-level model could not be reliably estimated. Sensitivity analyses using robust standard errors and a design effect correction were genuinely conducted, which is uncommon, but such corrections are considered unreliable below roughly forty clusters, a methodological judgment rather than a finding of the publication itself. | |
| Pre-randomization gap | Not tested |
| FindingThe methods place randomization after the baseline assessment and after the psychoeducation session. Yet part of the final between-arm gap is already present at that second measurement, the last one taken before the groups were formed: on the composite distress score, the gap widens from 0.47 point at baseline to 1.44 point after psychoeducation, then to 2.88 points by end of treatment. The authors test balance only at baseline. The attached protocol, by contrast, places randomization before the psychoeducation session: the appendices do not resolve the contradiction, they introduce it. | |
| Multiplicity | Not corrected |
| FindingAt least sixteen secondary comparisons, spread across four outcomes, two phases, and two contrast time points, without correction. With p values of 0.035 and 0.049, the question is not theoretical. | |
| Fit between findings and conclusion | Honest on substance |
| FindingThe negative primary outcome is stated in the body of the text and in the conclusion, the subgroups are labeled exploratory, and the discussion explicitly places the observed effects in the short-term register. The problem lies in presentation: the title erases the primary outcome, and the abstract attributes to the intervention phase effect sizes that are not, in fact, intervention-phase effect sizes. | |
| Scope | Narrow |
| FindingA single public high school, possible cross-class contamination acknowledged by the authors, a 15-day follow-up, a non-clinical population. No adverse event was reported, even though the protocol anticipated a possible transient increase in distress: worth noting, given that these are adolescents asked to write about their fears. | |
Level of evidence
What holds up: the trial shows no reduction in test anxiety beyond that produced by neutral writing, in this population and over this time horizon, even though it was powered to detect a d of 0.25 with 85% power. This negative result is clean, coming from a single, prespecified primary outcome in a properly randomized trial.
What holds up less: everything else. The differences on depression, anxiety, and sleep are small, uncorrected for multiplicity, measured by self-report among participants who knew which condition they were in, and already much reduced by 15 days on the composite score. Add to that the possibility that part of these differences predated randomization.
One distinction worth keeping: this result does not say that guided writing is useless. It says that it performs no better than writing twenty minutes a day about something else. That is not the same statement, and the active comparator is what makes the difference.
The colleague test
What an experienced colleague might say about this study in two minutes, between two consultations.
“The primary outcome is negative, so the trial is negative, full stop. That said, the protocol fits in three sentences, it costs next to nothing, and no adverse event was reported. A teenager panicking before a big exam, I could suggest this tomorrow, telling them it might help a little, not that it’s proven.”
Translated for practice: a low-cost option, with no adverse effect reported in this trial, and no demonstrated efficacy on test anxiety. That is not much, and it is already something when the alternative is offering nothing at all.
What you can do with this
- The three writing prompts are described precisely enough to reuse: three consecutive days, twenty minutes each. Day one, what the student thinks and feels about the exam. Day two, negative emotions and where they come from. Day three, positive emotions and what the student has learned about themself. In the trial, the prompts were preceded by a one-hour psychoeducation session and, each day, by a standardized video.
- Offer it for what it is: a possible aid, not demonstrated to reduce test anxiety. No adverse event was reported in this trial, though the protocol did anticipate a possible transient increase in distress while writing about negative emotions.
- Never present the 0.35 or 0.37 figures as effects of the intervention. They are between-group differences at a single time point, not effects of the intervention phase.
- Never cite the benefit on depression or sleep without stating, in the same breath, that the primary outcome was negative.
- Remember the configuration more than the result: when a title advertises three benefits, check which one was the primary outcome and what it actually showed.
- The course of action is set out in the NICE decision tree for generalised anxiety and panic disorder.
Frequently asked questions
Is this trial positive or negative?
Negative on its primary outcome, and the authors say so themselves. The secondary findings are small, and some of them are exploratory.
Why do the effect sizes in the abstract differ from those in the text?
Because they measure two different things. An interaction effect size measures a difference in trajectory between the two arms over time. A between-group difference at a single time point measures the gap at that instant. Here, the second is larger than the first by a factor ranging from 1.6 for insomnia to 2.5 for depression, and it is this second quantity that the abstract associates with the intervention phase.
Was the active comparator a sound choice?
Methodologically, yes: it controls for attention, time, and format. The cost is that nothing is known about how the intervention would compare with no intervention at all, and the authors acknowledge this among their limitations.
Does this transfer to students outside China?
The protocol does, it assumes nothing specific to the setting. The context does not: mandatory boarding, a single high-stakes exam, a distinct school culture. The underlying structure, expressing an emotion and then reappraising it, is the same one used in standard cognitive behavioral approaches.
Annotated bibliography
Source study. Luo Y, Fan J, Zang Y. Brief digital narrative intervention for adolescent depression, anxiety, and insomnia during academic stress: a cluster randomized controlled trial. BMC Medicine, 2026. Received May 31, 2025, accepted June 30, 2026. DOI 10.1186/s12916-026-05047-9. Trial registered with the Chinese Clinical Trial Registry under number ChiCTR2200058881. PMID 42400035, indexed in the PubMed registry on September 1, 2026. The document reviewed is the accepted manuscript, released by the journal ahead of final production: it carries no volume, issue, or article number, and the journal notes that the text will undergo further editing before final publication. The figures cited here were checked against the full text and against the document’s two appendices, the methodology and results appendix containing the study protocol, and the CONSORT 2025 checklist.
Editorial collections
Tags
Verified on September 1, 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on September 20, 2026, against the figures of the French version and against the source. How we verify what we publish
This analysis is intended for healthcare professionals. It does not constitute a prescribing recommendation and does not replace individual clinical judgment.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigor.
