Published on 14 September 2026
Does adding a standardised diagnostic tool to the child mental health pathway lead to more diagnoses?
In brief
One thousand two hundred and twenty-five children and young people aged 5 to 17 years, referred routinely to eight English National Health Service trusts for emotional difficulties, were randomised between usual care and usual care plus a structured diagnostic questionnaire completed online by the family or by the young person, whose summary report was uploaded to the service record. At twelve months, the proportion of children with a clinician-recorded diagnosis of an emotional disorder was 11% against 12%, adjusted risk ratio 0.94, 95% CI 0.70 to 1.28. The authors show no difference on any secondary outcome, and no analysis demonstrates cost-effectiveness. The qualitative evaluation offers an explanation: the report was not used consistently by clinicians during assessments, as had been intended. This is not a demonstrated absence of effect. It is an absence of evidence of effect.
The context
The idea is intuitive and widely shared: if emotional disorders in children go unrecognised, it must be because the clinician lacks information at the point of assessment. A structured questionnaire completed by the family beforehand should therefore increase the number of diagnoses made, shorten delays and improve triage. This system-level belief is rarely tested, because it looks self-evident. It has now been tested, at scale, under real service conditions.
The study at a glance
| Population | |
| Children and young people aged 5 to 17 years with emotional difficulties, referred routinely to child and adolescent mental health services in eight English National Health Service trusts, urban and rural. Excluded were emergency or urgent referrals requiring an expedited assessment, severe intellectual disability, and children already randomised in the trial. Randomisation ran from 27 August 2019 to 17 October 2021, a period straddling the pandemic. | |
| Intervention | |
| A structured diagnostic questionnaire (the development and well-being assessment) completed online by the family or by the young person after referral, with a summary report sent to participants and uploaded to the clinical record of the service, in addition to usual care. The report rested entirely on the algorithm-derived predictions, expressed in four bands running from close to average to very high. | |
| Comparator | |
| Usual care alone: review of the referral, clinical assessment if the referral is accepted, then offer and delivery of an intervention. | |
| Primary outcome | |
| A clinician-made diagnosis decision about the presence of an emotional disorder within twelve months of randomisation, collected from the clinical records of the service, with strict terminology: the suffix ‘disorder’ was required for the diagnoses that carry it. Uncertain wordings were referred for adjudication and classified as not constituting a diagnosis. A documented absence of diagnosis, like uncertainty, was coded as negative. | |
| Secondary outcomes | |
| Acceptance of the index referral and then of any referral, discharge, rereferral, confirmed diagnosis decision, diagnosis at eighteen months, time to diagnosis, to the decision to offer and to the start of treatment, participant-reported questionnaires at baseline and then at six and twelve months, economic measures of service use and quality of life. | |
| Design and numbers | |
| Pragmatic multicentre two-arm parallel-group randomised controlled trial, with a nested economic evaluation and a nested qualitative process evaluation. 1,225 randomised, 615 and 610 per arm, 610 against 609 analysed for the primary outcome. The sample size calculation required 1,210 participants to detect an absolute difference of ten percentage points, with 90% power and a two-sided 5% alpha, assuming 45% of diagnoses in the control arm and 10% of missing primary outcome data. |
Quality control
| Item checked | Judgement |
|---|---|
| Prospective registration | Verified |
| FindingTrial registered as ISRCTN15748675, with a favourable ethics committee opinion on 12 June 2019 and a first randomisation on 27 August 2019. The final protocol, version 4.1 of 1 August 2022, and the statistical analysis plan are public on the registry. The outcome definition and adjudication plan is dated 25 February 2020, during recruitment and well before database lock. A word-for-word comparison with the outcome as filed on the registry could not be made: the registry was not consulted. | |
| Allocation concealment | Sound |
| FindingMinimisation by algorithm, balanced on site, age band (5 to 10, 11 to 15, 16 to 17 years) and sex, minimising imbalance with 80% probability. The algorithm was created and concealed within a secure automated web system operated by the Nottingham clinical trials unit. | |
| Masking impossible in the field | Effective safeguard |
| FindingNeither participants, nor clinicians, nor the site researchers collecting source data were masked, and the authors count this among their limitations. Diagnoses identified in the records were recorded verbatim on the case report form, then submitted to an adjudication committee made up of the clinicians of the trial management group, blinded to allocation and to participant identifier; the trial data analysts worked blinded as well. | |
| Retention | Above 99% |
| FindingSix withdrawals of consent to access records out of 1,225, five in the intervention arm and one in the control arm, all of them before any emotional disorder diagnosis. The primary analysis covers every participant with observed outcome data, according to randomised allocation. That is the structural advantage of an outcome drawn from the records. | |
| Adherence to the intervention | 80% |
| Finding494 families out of 615, adherence being defined broadly as full or partial completion, a single module completed being enough. Among them, 332, or 67%, scored very high on at least one emotional disorder domain, most commonly depression and generalised anxiety. The authors themselves put the 20% who did not complete at the head of the possible explanations for the result. This is the most substantial reservation. | |
| Base rate assumption | 45% expected, 12% observed |
| FindingThe 45% assumption came from unpublished service data, which did not necessarily reflect the precise diagnostic terminology required in the trial. The authors note that combining firm and uncertain diagnoses gives 34 to 39% at twelve months and 39 to 43% at eighteen months, closer to what was expected at the outset. On the relative scale, the risk ratio interval, bounded at 0.70 and 1.28, excludes neither a 30% fall nor a 28% rise. The source publishes neither the absolute risk difference nor its interval. | |
| Nature of the outcome | Process outcome |
| FindingA diagnosis recorded in the notes is not a clinical outcome. The choice is coherent with the question asked, and the authors themselves set the next question among their future research priorities: whether receiving a clinical diagnosis makes a difference to health outcomes. | |
| Independence | Public funding |
| FindingIndependent research funded by the Health Technology Assessment programme of the National Institute for Health and Care Research, award number 16/96/09. No industry link is declared. Institutional links with the funder do exist: the chief investigator is an NIHR senior investigator and sits on its HTA prioritisation committee, and another author sat on the clinical evaluation and trials committee of the same programme. Since the result is negative, the hypothesis of complaisance does not hold. | |
The findings
| Outcome | Reported value |
|---|---|
| Primary outcome at 12 months | 68/610 against 72/609, adjusted RR 0.94 [0.70 to 1.28], p = 0.71 |
| ReadingThe authors report no evidence of a difference between groups. The sample size calculation targeted an absolute difference of ten points under the assumption of 45% of diagnoses in the control arm, whereas the observed rate is 12%. The source publishes neither the absolute risk difference nor its interval: from this source alone, one cannot say what size of absolute benefit is ruled out. | |
| Diagnosis at 18 months, secondary outcome | 14% against 15% |
| ReadingThe result does not move with the additional follow-up. The document consulted gives no quantified effect estimate for this outcome. | |
| Referral acceptance | 277 (45%) against 262 (43%), RR 1.06 [0.94 to 1.19] |
| ReadingFor any referral by eighteen months, 374 (61%) against 352 (58%), risk ratio 1.06, interval 0.97 to 1.16. The authors themselves note a suggestion of a small influence, of 3 to 6%, on the likelihood of acceptance of the index referral or of any referral by twelve or eighteen months. This is the only place where they leave the hypothesis of an effect open. | |
| Other secondary outcomes | No difference shown |
| ReadingThe authors report no evidence of any differences between groups across all the other secondary outcomes, whether these come from the records or from the participant-reported questionnaires. The document consulted does not detail the corresponding estimates. | |
| Economic evaluation | No analysis demonstrates cost-effectiveness |
| ReadingPrimary perspective of the health service and personal social services, primary outcome measure the quality-adjusted life-year, secondary analysis from a societal perspective including productivity losses and out-of-pocket expenses of families, between-group differences estimated by seemingly unrelated regressions on multiply imputed data. The authors write that they found “no evidence to suggest that the intervention impacted health service utilisation, broader societal costs or quality-of-life outcomes” for the children or their parents and carers. The quantified results do not appear in the document consulted: it refers for these to the appendix of the primary publication. | |
Critical appraisal
| Domain | Judgement |
|---|---|
| Randomisation and allocation | Low risk |
| FindingCentral minimisation, automated allocation through a secure web system, and baseline characteristics described as well balanced between the groups, for the children as for the parents. | |
| Deviations from the intervention | Some concerns |
| FindingMasking impossible in the field, and adherence of 80% in the broad sense, completion of a single module being enough to count as adherent. In a trial that shows no difference, this last figure is the main admissible objection. | |
| Missing data | Low risk |
| FindingRetention above 99% on the primary outcome, six withdrawals out of 1,225. | |
| Outcome measurement | Concerns mitigated |
| FindingThe researchers extracting the records were not blinded, but verbatim recording followed by adjudication blinded to allocation and to participant identifier neutralises most of the risk. The authors add that any data inaccuracies would need to be substantial for any true effect to be negated. | |
| Selective reporting | Low risk |
| FindingRegistration, public protocol and statistical analysis plan, adjudication plan for the primary outcome dated and prior to database lock, negative conclusion stated without an escape route. The document consulted reports no subgroup analysis: that is an absence of analysis in this source, not an absence of signal. | |
Level of evidence
Oxford level of evidence 1b: a single randomised controlled trial, multicentre, large, at low risk of bias on randomisation, on missing data and on outcome measurement. Confidence is high that this tool, used in this way, in this health system, did not show an increase in the number of diagnoses recorded in the notes, nor an economic benefit. It is more reserved on how far the result carries outside the system that produced it, and on what should be concluded about a set-up in which clinician buy-in was actually secured. The authors bound their own conclusion: they write that they found “no evidence that completion of the development and well-being assessment aided the detection of emotional disorders in this study”, and that using it in this way “cannot be recommended for clinical practice”. Those bounds must be kept in any restatement of the result.
The colleague test
What an experienced colleague would say if you put this study to them in two minutes, between two consultations.
“ A British trial in 1,225 children finds no evidence that sending families a structured diagnostic questionnaire before the assessment leads to more diagnoses, shorter delays or lower costs. The qualitative arm suggests why: few of the clinicians interviewed had used the report in the consultation. ”
What this means in practice: before funding another data-collection tool, ask who will read the document it produces, at what point in the decision, and what will change in the consultation. The limiting factor did not appear to be the information available, but its entry into the clinical decision.
What you can do with this on Monday morning.
- Do not expect a structured questionnaire sent to families to increase the detection of emotional disorders in children on its own.
- When a data-collection tool is deployed in a service, plan from the outset how the document will be read, when it will be used and where it sits in the decision. Without that, the result of this trial is the result to expect.
- Keep the wording accurate: there is no evidence of an effect. Writing that there is no effect claims more than the trial shows, and the source does not publish the absolute difference that would tell you what size of benefit is ruled out.
- Use this trial as a methodological reference when an organisational measure is proposed without evaluation: one large, rigorous negative trial is worth more than a shared intuition.
- The course of action is set out in the NICE decision tree for generalised anxiety and panic disorder.
Frequently asked questions
Is the tool at fault, or the way it was used?
The trial tests one tool used in one precise way. The qualitative evaluation identifies several converging obstacles: workload, difficulty finding the report inside the electronic record, the delay between the report being generated and the first assessment, and clinician reluctance towards diagnostic terminology. That points towards the use rather than the tool, but neither hypothesis is demonstrated separately.
Do these findings apply outside the English health service?
The trial was run in eight English National Health Service trusts, whose organisation differs from that of child and adolescent services in other countries. Direct transferability was not studied, and the authors explicitly bound their conclusion to this system and to this delivery format. The operational message does travel, at the moment when services digitise their intake procedures.
Was the trial adequately powered?
It exceeded its recruitment target, 1,225 randomised against 1,210 planned, with power calculated at 90% for an absolute difference of ten points. The assumption of 45% of diagnoses in the control arm turned out to be far above the 12% observed, which widens the uncertainty on the relative scale. Since the source does not publish the absolute risk difference, one cannot deduce from it what size of absolute benefit remains compatible with the data.
Does the result vary with age or sex?
The document consulted reports no subgroup analysis. That is an absence of analysis in this source, not an absence of difference: nothing can be concluded from it either way.
Why does a negative trial of this size get so little attention?
Because it offers nothing new to do. That is precisely why it deserves to be read: it saves the effort of deploying a measure whose benefit was looked for and not found.
Annotated bibliography
Source report. Sayal K, Wyatt L, Thomson L, Holt G, Ewart C, Bhardwaj A, Dubicka B, Marshall T, Gledhill J, Lang A, Sprange K, Partlett C, Newman K, Moody S, Bould H, Upton C, Keane M, Cox E, James M, Montgomery A. Clinical and cost-effectiveness of a standardised diagnostic assessment for children and adolescents with emotional difficulties: the STADIA multi-centre RCT. Health Technology Assessment 2025;29(61):1-34. DOI 10.3310/GJKS0519. PMID 41239894. Synopsis of the Health Technology Assessment programme. This is the reference document for this analysis. It reproduces, with permission, material from the two publications below, and refers explicitly to the primary publication for the detailed tables and the quantified economic results, which it does not contain.
Primary trial publication. Sayal K, Wyatt L, Partlett C, Ewart C, Bhardwaj A, Dubicka B, et al. The clinical and cost effectiveness of a STAndardised DIagnostic Assessment for children and adolescents with emotional difficulties (STADIA): multi centre randomised controlled trial. Journal of Child Psychology and Psychiatry 2025;66:805-20. DOI 10.1111/jcpp.14090. PMID 39775729. Reference cited in the source report, which supplies its digital object identifier and its PubMed identifier. It was not consulted for this analysis: no figure reported here comes from it.
Qualitative process evaluation. Thomson L, Newman K, Ewart C, Bhardwaj A, Dubicka B, Marshall T, et al. Barriers and facilitators to using standardised diagnostic assessments in child and adolescent mental health services: a qualitative process evaluation of the STADIA randomised controlled trial. European Child and Adolescent Psychiatry 2025;34:2763-77. DOI 10.1007/s00787-025-02678-w. PMID 40100401. Source of the qualitative strand, summarised in the report and in its appendix 3. Not consulted directly.
Published protocol. Day F, Wyatt L, Bhardwaj A, Dubicka B, Ewart C, Gledhill J, et al. STAndardised DIagnostic Assessment for CYP with emotional difficulties (STADIA): protocol for a multicentre randomised controlled trial. BMJ Open 2022;12:e053043. Trial registered as ISRCTN15748675. The source report gives neither a digital object identifier nor a PubMed identifier for this protocol: they are left out here rather than guessed.
What was consulted. Verification carried out on 13 August 2026. The 58 pages of the Health Technology Assessment programme report were read in full: structured abstract, plain language summary, introduction and objectives, methods of data collection and analysis, economic analysis, summary of the results of the main trial and of the qualitative evaluation, discussion and interpretation, patient and public involvement, equality and diversity, implications for decision-makers, practice and research recommendations, conclusions, additional information (contributions, disclosure of interests, registration, funding), reference list, then the three appendices: protocol amendments, schedule of assessments, eligible diagnoses and secondary outcomes; summary of the main trial; summary of the qualitative process evaluation. The four supplementary material documents were also consulted in full: referral screening form, consent and participation table, outcome definition and adjudication plan of 25 February 2020, and the template of the report given to families. The bibliographic metadata, volume, issue, pagination, digital object identifier and PubMed identifier, were established from the publisher and PubMed records. One point to flag: the figures of the economic evaluation, the absolute risk difference and the subgroup analyses appear in none of the documents consulted, the report referring for these to the appendix of the primary publication, which was not consulted; these elements were therefore left out rather than reproduced. This analysis underwent an independent double reading.
