Published on 16 September 2026

Analysis · All conditions · Methodology

Psychiatric guidelines: 11.6% rated high level of evidence. Should we distrust them for that?

▤ Dossier BMJ Mental Health · 2026; 29 (1): e302194 · Rømer et al. DOI 10.1136/bmjment-2025-302194 PMID 41806975 Scientific 70 Editorial 89

In brief

A systematic review published in BMJ Mental Health went back over 545 recommendations, extracted from 24 documents published between 2014 and 2024 by four issuing bodies: the American Psychiatric Association, the European Psychiatric Association, the World Federation of Societies for Biological Psychiatry and the World Health Organization. For each recommendation, the authors recorded two distinct things: the level of evidence as the guideline authors themselves had assigned it, harmonised to the GRADE vocabulary, and the class of the highest-level study actually cited in support. The two measures diverge, and that is the central result: 63 recommendations, or 11.6%, are rated high level by their own authors, whereas 241, or 44.2%, cite at least two randomised trials or a meta-analysis of such trials. In other words, guideline authors heavily downgrade the evidence from the trials available to them. No significant trend emerges over the ten-year period (p = 0.07). This work describes the state of the evidence cited, it does not judge the clinical soundness of the recommendations, and the authors say so themselves.

The context

In the consulting room, a practice guideline works as a trusted shortcut. It condenses a body of literature that no clinician can read in full, and it provides support when a decision is questioned, by the patient, by a colleague or by a third party. Its strength lies in that economy: one follows the line without having to demonstrate again what underpins it.

The price of this shortcut is that it makes the solidity of what has been condensed invisible. Two recommendations written in the same document, with the same typographical authority, may rest, the one on several concordant randomised trials, the other on a case series and expert consensus. The hurried reader does not tell them apart, all the more so as not every body displays its level of evidence with the same clarity, or on the same scale.

The question this work asks is therefore not clinical but epistemic: when we apply a recommendation, what exactly are we standing on? It is worth asking because in most cases the answer does not change what we do, but it changes how we defend and discuss it in the cases where uncertainty matters.

What GRADE measures

Certainty of evidence and strength of recommendation are two separate axes

A methodological reminder, outside the scope of this study. GRADE does not rate the quality of a recommendation. It rates the certainty that can be placed in the effect estimate supporting it, on four levels: high, moderate, low, very low. That certainty is downgraded for risk of bias in the studies, imprecision of the estimates, heterogeneity, indirectness of the comparisons and risk of publication bias.

The strength of the recommendation is a second judgement, which takes in the balance of benefits and harms, patients’ values and preferences, feasibility and cost. A strong recommendation backed by low certainty is entirely consistent within the GRADE framework: it is the typical case of a course of action that no one will ever test in a randomised trial, either because such a trial would be considered unethical or because the event to be prevented is too rare to be studied with sufficient power.

Two clarifications are needed about how this scale is used here. First, the review authors did not rate the recommendations themselves: they recorded the rating each body had already assigned, then translated the different bodies’ heterogeneous scales into the common vocabulary of GRADE. Second, the other axis was indeed measured: the declared strength of each recommendation, strong or weak, was recorded and cross-tabulated with the level of evidence. The cross-tabulation speaks for itself. Among strong recommendations, 22.9% rest on a high level and 45.0% on a moderate level, against 0.4% and 15.5% among weak recommendations. And 44 of the 249 strong recommendations, or 17.7%, cite only cases or expert opinion in support.

The study at a glance

Systematic review with quantitative synthesis, pre-registered protocol
Population
CorpusGuidelines published from 2014 to 2024 by four pre-selected bodies: the American Psychiatric Association, the European Psychiatric Association, the World Federation of Societies for Biological Psychiatry and the World Health Organization. Search of the four bodies’ websites on 15 July 2024, updated on 22 June 2025. No condition is targeted in particular: the object is the guideline literature itself, with secondary stratification by diagnostic group.
Intervention
ProcedureFor each recommendation, recording of the declared strength, strong or weak, and of the level of evidence assigned by the guideline authors themselves, harmonised to the GRADE vocabulary. Separate recording of the highest-level study cited in support, classified according to the ACC/AHA nomenclature, with its blinding and its comparator. Extraction by one reviewer, full validation by a second, arbitration by a third in case of disagreement. No inter-rater agreement statistic is reported in the publication.
Comparison
AxesBreakdown by issuing body, by type of recommendation, by diagnostic group and by year of publication. Pre-specified comparison with equivalent work conducted in other specialties. Two sensitivity analyses, one restricted to the most rigorous documents, the other to the last five years.
Outcome
MeasureTwo distinct measures. Proportion of recommendations classified as high, moderate, low or very low level according to the guideline authors’ appraisal. Proportion citing class A, B or C evidence, class A denoting at least two randomised trials or a meta-analysis of such trials, class B a single trial or observational studies, class C cases or expert opinion.
Sample analysed
Volume82 documents screened, 24 included, 545 recommendations extracted. The main reason for exclusion is the lack of any update for more than ten years, which rules out 29 documents, or 35% of the documents screened.
Design and level of evidence
DesignSystematic review of guidelines with quantitative synthesis, a piece of meta-research. No CEBM level is applicable or mentioned: the CEBM scale addresses clinical questions, not a body of guidelines. The quality of the 24 included documents is appraised with AGREE II, domain 3, “Rigour of Development”, by a single reviewer; 6 documents out of 24 reach a score of at least 70%.

Quality control

CheckpointJudgement
Scope of the corpusPre-selected, not exhaustive
Finding545 recommendations over ten years, from the four most influential issuing bodies. The corpus is nonetheless not exhaustive, and the authors say so: only documents in English that explicitly grade the level of evidence and cite the evidence recommendation by recommendation could be retained. NICE is excluded for that reason. Guidelines from low- and middle-income countries were not accessed, because of the language barrier. The authors write that their results may not generalise to other regions or to other subfields.
Prior registrationProtocol registered before extraction
FindingThe protocol was registered on the Open Science Framework before data extraction began, and the data and analysis code are deposited there. The comparison with other specialties was pre-specified. This is the clearest methodological strength of the work.
Conduct of the reviewPRISMA stated
FindingThe flow diagram and the PRISMA checklist are stated to be in the supplementary material, which we were unable to consult. We therefore cannot check the completeness of that checklist ourselves.
Data extractionA single extractor
FindingExtraction was carried out by a single reviewer, all extracted data being validated by a second reviewer, with a third author consulted in case of disagreement. No inter-rater agreement statistic is reported, so the share of judgement involved in classifying recommendations by type and by diagnostic group cannot be quantified. The AGREE II appraisal, for its part, rests on a single reviewer with no stated validation.
Origin of the level of evidenceGuideline authors’ rating, not reviewers’
FindingA point not to be confused. The level of evidence analysed was not assigned after the fact by the Danish team: it is the one each body had itself published, a condition for the document’s inclusion. The authors’ work consisted of translating heterogeneous scales into a common vocabulary. The room for interpretation therefore lies in that translation, not in a substitute rating. In return, the authors point out that the guideline authors’ appraisal can be influenced by their pre-existing beliefs and conflicts of interest, which is precisely why they added the independent recording of the studies cited.
Funding and competing interestsNo links with the bodies assessed
FindingA team from the Copenhagen Research Center for Biological and Precision Psychiatry and the University of Copenhagen, with no links to the four bodies analysed. No specific funding was declared for this work. Outside the submitted work, the senior author declares being principal investigator on grants from the European Research Council, the Capital Region of Denmark, the Mental Health Services of that same region, the Lundbeck Foundation and the Novo Nordisk Foundation; the first author declares a grant from the Novo Nordisk Foundation. The last two are foundations linked to pharmaceutical groups: the link is indirect and unrelated to the subject here, but the reader should know about it.

The findings

11.6%
of the 545 recommendations analysed are rated high level of evidence by their own authors, that is 63 recommendations. At the same time, 241 recommendations, or 44.2%, cite at least two randomised trials or a meta-analysis of such trials. This gap is the central result of the work.
OutcomeReported value
Distribution by level of evidence, whole corpusHigh 63 (11.6%), moderate 165 (30.3%), low 128 (23.5%), very low 189 (34.7%)
ReadingSlightly more than a third of the corpus falls at the lowest level of the scale. These four values are those assigned by the guideline authors, translated into the GRADE vocabulary.
Class of the best study citedClass A 241 (44.2%), class B 167 (30.6%), class C 137 (25.1%)
ReadingClass A breaks down into 155 recommendations (28.4%) citing a meta-analysis of randomised trials and 86 (15.8%) citing at least two trials. This is where the gap lies: 44.2% of the corpus relies on what the nomenclature regards as the strongest evidence, but only 11.6% derive a high level of certainty from it. The authors see this as guideline authors downgrading the evidence from the trials available to them.
Share of very low level, by bodyAPA 65 (59.6%), WFSBP 90 (36.7%), WHO 12 (19.7%), EPA 22 (16.9%)
ReadingThe American Psychiatric Association shows the highest proportion of very low levels. No statistical test compares the bodies with one another, so this ranking is descriptive. The indicator the authors highlight is, moreover, the other end of the scale: the share of high level ranges from 20.0% for the European Psychiatric Association to 1.6% for the World Health Organization.
Trend from 2014 to 2024No significant trend, p = 0.07
ReadingThe proportion of recommendations at high or moderate level shows no significant trend over the period, a logistic regression on year of publication giving p = 0.07. Neither an improvement nor a deterioration is demonstrated. A sensitivity analysis restricted to documents published from 2019 to 2024 finds a similar proportion, 30 recommendations out of 311, or 9.6%.
Differences by type of recommendationFrom 14.6% for pharmacological treatments to 0% for somatic assessment
Reading41 of the 281 recommendations on pharmacotherapy are at high level, against 1 of 46 for organisation of care and 0 of 13 for somatic assessment of patients with severe mental disorders. Some areas count fewer than ten recommendations in total, patient information, prevention, discontinuation of treatment, treatment goals, and support no fine-grained interpretation.
Post hoc analysis of the best recommendations8 meta-analyses out of 25, or 1.5% of the corpus
Reading25 recommendations combine a high level with citation of at least one meta-analysis. Among the meta-analyses cited, only 8 provide an effect estimate based exclusively on double-blind trials with an adequate comparator, that is 1.5% of the 545 recommendations, and all concern pharmacotherapy. Conversely, 11 of 25 rest only on trials that were not double-blind and had no comparable control. A post hoc analysis, to be treated as hypothesis-generating.
Statistical precisionCounts and proportions, no confidence intervals
ReadingThe publication reports counts and percentages. No confidence interval is provided, and the only inferential test is the logistic regression on year of publication. It is therefore not possible to say whether the differences between bodies or between types of recommendation exceed sampling fluctuation.

Critical appraisal

DomainRisk
Harmonisation of scalesTranslation not quantified
FindingThis is the weightiest point. The four bodies do not use the same scale, and the authors had to bring those scales down to a common vocabulary. This translation involves a share of judgement that is not quantified: no inter-rater agreement statistic accompanies the work, and the correspondence tables are in the supplementary material, which we were unable to consult. Any comparison between bodies depends directly on the validity of this translation.
What a high level really measuresA declared rating, not audited
FindingThe figure of 11.6% does not say that 11.6% of recommendations are well supported. It says that their authors declared them so. The authors point out that this appraisal can be influenced by the guideline authors’ pre-existing beliefs and conflicts of interest, which cuts both ways: a body that is scrupulous in displaying its uncertainties penalises itself in this kind of count. That is exactly why the independent recording of the studies cited is the useful counterweight in this work.
A single study retained per recommendationAcknowledged by the authors
FindingThe ACC/AHA classification retains only the highest-level study cited. The authors acknowledge that this choice ignores the bulk of the evidence actually drawn on, which may lie elsewhere, in numerous observational studies for example. Nor did they assess whether the studies cited actually supported the recommendation, or the quality of each of them.
Selection of documentsRestricted by design
FindingTo enter the corpus, a document had to grade its level of evidence recommendation by recommendation. This criterion is essential to the project, but it rules out influential texts, NICE first and foremost, and it selects the bodies that already display their uncertainties. English language as a criterion rules out all non-English-language national guidelines.
Funding and competing interestsDeclared, no links with those assessed
FindingNo specific funding for this work, no links with the bodies analysed, a complete ICMJE declaration mentioning unrelated research grants, including two foundations linked to the pharmaceutical industry. An expected configuration, properly documented.
Fit between claim and evidenceCalibrated conclusions
FindingThe authors write explicitly that concluding that 11.6% of recommendations rest on high-level evidence does not mean that 11.6% of clinical decisions rest on high-level evidence, since guidelines are not always followed and recommendations do not all have equal impact. They also point out that the comparisons with other specialties, where values range from 5% to 23%, rest on different methodologies and should not be over-interpreted. This restraint deserves emphasis, because that is exactly where media coverage usually goes astray.
Transfer to national guideline corporaNot assessed
FindingNational guidelines that clinicians use day to day, for example those of the French Haute Autorité de santé for psychiatrists practising in France, are not part of the corpus analysed. This work therefore says nothing about their level of evidence, in either direction, and extrapolating would be an opinion here, not a result.
What the design does not allowNo causal inference
FindingThe work is descriptive. The authors write that the reason for the downgrading of the evidence is not directly addressed in their study. They put forward plausible explanations, indirect trials, clinically irrelevant questions, conflicting results, low power, unblinding, missing data, irrelevant endpoints, publication bias, but these explanations are hypotheses, not results.

Level of evidence

Scientific70/100
Editorial89/100

What is demonstrated, in the sense of a reproducible count on a defined corpus with a protocol registered before extraction: the share of recommendations declared high level by their authors is 11.6%, the share citing at least two randomised trials is 44.2%, and the gap between these two values is massive. The count is independent of the bodies assessed, published in a peer-reviewed journal, with data and code publicly deposited.

What is suggested, and not demonstrated: that this gap reflects guideline authors downgrading the evidence rather than an artefact of translating between heterogeneous scales. The distinction is not academic, since no reproducibility statistic accompanies this translation. Also suggested, without being tested, is the relative ranking of the bodies: no test compares the publishers, no confidence interval is provided, and the fact that a body shows the highest share of very low levels may reflect the rigour with which it displays them as much as the weakness of its foundations.

What is not established, and must not be read into this work: any change over the period. No significant trend emerges, p = 0.07. Neither progress nor decline.

That leaves what belongs to expert opinion, and must be identified as such. The authors recommend a shared approach to evidence appraisal across bodies, a commitment to updating documents within years rather than decades, and priority funding for trials in under-evidenced areas, organisation of care and somatic comorbidities. These are defensible positions, they are not results of the study.

The colleague test

What an experienced colleague would say if you put this study to them in two minutes, between two consultations.

“ On Monday morning, I read guidelines differently. The most unsettling thing is not that 11.6% are high level, it is that almost four times as many cite randomised trials. That means the guideline authors themselves do not put much faith in the trials they cite. It does not mean I am going to stop following guidelines, it means I know where I am on firm ground and where I am on soft ground. And frankly, on some questions, nobody will ever run the missing trial. ”

What this means in practice: the level of evidence is not a grade given to the recommendation, it is information about the solidity of what supports it. It serves to calibrate what we say, not to sort good recommendations from bad ones.

What you can do with this

  • Get into the habit of reading, alongside each recommendation, the stated level of evidence and not only the strength of the wording. In most texts this information exists, it is simply relegated to small print or an appendix.
  • Systematically separate two questions: how certain is the evidence, and how strong is the recommendation. The corpus analysed shows that 17.7% of strong recommendations cite only cases or expert opinion in support.
  • Look at what is cited as well. A recommendation that refers to a meta-analysis of randomised trials does not have the same backing as one that refers to a case series, even when both carry the same colour in the summary table.
  • When faced with a low-certainty recommendation, make that uncertainty explicit in shared decision-making and record it in the notes. This is the situation in which alternatives are genuinely discussed with the patient, and in which the preference the patient expresses weighs most.
  • Be wary of poorly documented areas rather than of bodies. Somatic assessment of patients with severe mental disorders has no high-level recommendation among the thirteen identified, and organisation of care only one out of forty-six.
  • Do not transfer this result to national guidelines that are not in the corpus analysed, for example those of the French Haute Autorité de santé. The reading habit transfers, the figure does not.
  • Resist the catchphrase that will circulate: a low level of evidence does not make a recommendation bad. It indicates that we are moving forward with less documented margin for error, which is not the same thing.

Frequently asked questions

Does a low GRADE level of evidence mean a guideline recommendation is wrong?

No, and this is the main confusion to avoid. GRADE describes the certainty of the effect estimate, not the soundness of the recommended course of action. A recommendation can be strong, sound and universally followed while resting on very low certainty, simply because the randomised trial that would demonstrate it will never be carried out, for ethical reasons or for reasons of power. The authors also point out that a trial is not always necessary or relevant, notably when the intervention is a basic human right such as housing or access to healthcare. A low level is information about the state of the literature cited, not a verdict on the recommendation.

Who assigned the levels of evidence analysed in this review?

The guideline authors themselves. This point is often misunderstood. A document had to grade each of its recommendations explicitly in order to enter the corpus. The review authors recorded those ratings, then translated them into the GRADE vocabulary because the four bodies do not use the same scale. Independently, they added a record of the highest-level study cited in support of each recommendation. It is these two measures that diverge.

Is class A evidence the same as a high GRADE level of evidence?

No, and that is precisely the finding of the study. In the nomenclature used here, borrowed from the American cardiology societies, class A denotes a level of evidence, that of at least two randomised trials or a meta-analysis of such trials. It does not denote the strength of the recommendation, which in that system has a separate numbering. Yet 44.2% of recommendations cite class A evidence while only 11.6% are declared high level. Reading class A as a guarantee of high certainty is a translation error between two scales.

Do these findings apply to national psychiatry guidelines?

The corpus analysed is limited to four bodies: the American Psychiatric Association, the European Psychiatric Association, the World Federation of Societies for Biological Psychiatry and the World Health Organization. National guidelines, for example those of the French Haute Autorité de santé, are not included. We therefore do not know what the same analysis would show when applied to such a corpus, and asserting it in either direction would be conjecture.

Does this mean clinicians should stop following guidelines?

No, and the authors themselves do not conclude so. For the individual clinician, a guideline remains the best available synthesis of a literature that cannot be read in full. This work is an invitation to read it together with its level of evidence, not to do without it. The alternative to a weakly supported recommendation is not the absence of a recommendation, it is individual judgement, whose certainty is no better documented.

What is the main limitation of this review?

It lies in the harmonisation of scales. The four bodies do not use the same rating scale, and the authors had to bring them down to a common vocabulary in order to compare them. This involves a share of judgement that is quantified nowhere, for want of an inter-rater agreement statistic, and extraction moreover relied on a single reviewer, validated by a second. On top of that, the record of the study cited retains only the highest-level reference, which ignores the bulk of the evidence actually drawn on, a limitation the authors explicitly acknowledge.

Annotated bibliography

Rømer TB, Andersson SN, Benros ME. Levels of evidence supporting American, European and international guidelines in psychiatry, 2014-2024: a systematic review with quantitative synthesis. BMJ Mental Health, 2026, volume 29, issue 1, article e302194. DOI 10.1136/bmjment-2025-302194. PMID 41806975.
A systematic review with quantitative synthesis of 545 recommendations extracted from 24 documents published from 2014 to 2024 by four international bodies, recording the level of evidence declared by the guideline authors, harmonised to the GRADE vocabulary, and independently recording the highest-level study cited. Pre-registered protocol, data and code deposited on the Open Science Framework. Contribution: an independent count of something rarely quantified, and evidence of a gap between the evidence cited and the certainty declared. Limitation: the harmonisation of heterogeneous scales comes with no reproducibility statistic, and a single reviewer carried out the extraction.

What was consulted. Figures, method and references verified by independent double reading against the full text of the published version, consulted on 13 August 2026, and against the PubMed record for the identifiers, affiliations, funding statement and competing interests statement. The publication’s supplementary material, a single file containing tables S1 to S11 and figures S1 and S2, could not be consulted: it was not provided and could not be retrieved from the publisher. Items that exist only in that material, notably the correspondence tables between the four bodies’ scales, the PRISMA flow diagram and the details of the sensitivity analyses, are therefore reported here only on the basis of what the main text says about them, and this is flagged each time. No confidence interval, no inter-rater agreement statistic and no CEBM level appear in the publication: they do not appear in this analysis either.

Editorial collections

Topics

Verified on 13 August 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on 16 September 2026, against the figures of the French version and against the source. How we verify what we publish.
Content published by Psychiatry Evidence Base is produced according to the principles of evidence-based medicine. Every analysis rests on an independent critical reading of the scientific literature and aims to help health professionals interpret it. The information presented replaces neither official guidelines, nor clinical reasoning, nor individualised care. Medicine evolves continuously, and some data may change as new scientific evidence appears.
Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigour.

Report an error in this analysis

Follow Dr Stroescu on LinkedIn, for the review every Saturday