Resource centre · Understand
Our methodology
How a scientific paper is selected, appraised, then published or set aside on Psychiatry Evidence Base. The principles are public, the thresholds stay internal, and this page explains why.
In one sentence. A paper is never selected because it is recent or because it will provoke a reaction: it is selected because it is solid, and how it is published is decided afterwards.
01 Two axes, never mixed
Every publication is appraised twice, separately, on two questions that have nothing to do with each other.
The scientific axis asks: can what this paper claims be trusted? It looks at the level of evidence, the risk of bias, statistical precision, the fit between the size of the conclusions and the strength of the design, independence and transparency.
The editorial axis asks: does this paper deserve to be published, and in what form? It looks at clinical reach, immediate applicability in French office practice, the importance of understanding it even where transposition is not yet possible, teaching potential, lasting value, and the angle that is PEB’s own.
Plainly. The first axis decides whether the paper has the right to exist on the site. The second decides only the place it occupies there. If the two were added together, a fascinating subject could rescue a badly done study. That is what we are avoiding.
02 A floor that nothing redeems
The editorial score never compensates for the scientific score. Below a certain level of solidity, a paper stays in the documentary base and is not published. However fascinating it is.
Three situations trigger this cap: overall scientific solidity that is insufficient, a weakness on the level of evidence or on the risk of bias, and conclusions larger than the design can support. The third is the most common, and the hardest to see. The study is sound. It is its conclusions that overflow.
03 Facts before judgement
Before any overall appraisal, a list of objectifiable items is filled in: registration or pre-registration of the study, randomisation and masking according to design, sample size or number of included studies, declaration of conflicts of interest, funding source, data availability, existence of a replication, pre-specified primary outcome, clinical and not merely statistical effect.
Plainly. This step is not there for show. It neutralises the halo effect, the reflex that makes us judge the whole of a paper favourably once its first page has convinced us. First tick what can be checked, then judge.
04 The instrument follows the document
A randomised trial is not appraised with the instrument made for a guideline. The type of document is identified first, and it determines the instrument.
| Type of document | Risk-of-bias instrument | Level-of-evidence marker |
|---|---|---|
| Randomised controlled trial | RoB 2 | Oxford levels 1b or 2b |
| Meta-analysis, systematic review | AMSTAR-2, ROBIS | Level 1a, with heterogeneity and publication bias examined |
| Observational study | ROBINS-I | Levels 2b to 4 |
| Practice guideline | AGREE II | Appraised on development quality, not on risk of bias |
Two common confusions are set aside at the outset. GRADE appraises the certainty of a body of evidence, not a single paper: for one publication, the design hierarchy is the reference. And CONSORT, PRISMA or STROBE are reporting guides, not risk-of-bias instruments: a perfectly reported paper can be methodologically poor.
05 Steps, not false precision
Every criterion is rated on one of five steps only: absent or disqualifying, weak, acceptable, solid, exemplary. No continuous rating.
Plainly. Nobody can honestly tell a paper rated 18 from a paper rated 19. That precision is decorative, and it gives a subjective judgement the appearance of a measurement. Five anchored steps produce a less pretty and more reproducible result.
06 A human at the borders
A paper whose score places it right on the line between two decisions never tips over automatically. It triggers a full re-reading, and the decision is taken again by hand.
It is at the borders that an instrument of this kind is most fragile, because one point of difference changes a conclusion there while reflecting no real difference. So a reader is put there, not a decimal. Three operations in particular are never carried out without an explicit human ruling: moving a paper up into a higher category, changing a score to justify a decision already taken, and publishing a paper that has crossed below the scientific floor.
07 The colleague test
No numerical instrument captures everything. The last filter is therefore not a score, it is a question, asked out loud before any publication.
The question
Would a demanding peer I respect wince at seeing me pass this on?
If the answer is yes, the paper does not go through, whatever its score. That question does work the grids cannot do. A publication can tick every box and still be awkward: because the message it will produce once summarised will be stronger than what it demonstrates, because its funding makes its conclusion predictable before reading, because it arrives at a moment when it will be picked up for what it does not say.
Plainly. The threshold excludes what is fragile and what is awkward, not everything that is imperfect. No paper is perfect. The question is not “is this beyond reproach”, it is “am I ready to answer for it in front of someone who knows the field”.
It is the least formalisable part of the method, and probably the most useful. At the very end of a highly structured appraisal chain, it brings back the one thing a structure does not replace: the responsibility of whoever signs.
08 What this method is not
Three clarifications, out of honesty and because the question is a fair one.
It is not peer review
PEB is written by a single author, a practising psychiatrist. The instrument described here does not replace the peer review of a scientific journal: it makes the judgement of a single reader explicit, reproducible and checkable. It is less than an editorial board, and it is a great deal more than a selection made along with the news.
It is not a source of guidelines
The analyses published here comment on scientific publications and on existing guidelines. They replace neither those texts nor personalised medical advice. The reference texts in force remain the ones to consult.
The thresholds are not published, and that is deliberate
The weightings, the floor values and the decision matrix stay internal. Two reasons. A published value becomes wrong at the next version of the instrument, without anyone noticing, and a reference file that ages in silence is worse than no file at all. And a scale that is entirely public ends up being optimised against itself. The principles, on the other hand, do not change, and it is the principles that let you judge the method.
09 How we verify what we publish
A PEB page puts forward sentences, figures and dates. Here is what we do before publishing them, and what we cannot promise.
Documents are frozen before they are read
When we work on a guideline, we download all of its parts and compute a digital fingerprint for each, a sort of barcode of the file. If the file changes later, the fingerprint changes and we see it. That avoids citing a version we could no longer find.
A quotation is taken from the document, never from memory
Every sentence in quotation marks is taken word for word, with its chapter, its page and its date of approval. When we cut a sentence that is too long, we mark it with brackets. We do not reproduce the bullet lists of our sources: what we draw from them is announced as a summary, and it is written by us.
Figures are recounted, not copied
How many statements carry a level of evidence, how many times a phrase recurs in a text, how many patients in a study: we count ourselves, on the full text, by programme. A figure that cannot be found in the source does not go out. And if a table expects a value the source does not give, we write that the source does not give it. That is already information, often the most useful kind.
Beware of pages that are images
Many publications set their diagrams as images. A search inside the file then finds nothing, and what matters is missed. It happened on the ADHD dossier of the French health authority: the dose that contradicted the rest of the dossier sat on a page set as an image. When a document contains diagrams, we look at them as images, not as text.
An absence is the most fragile claim of all
Writing “this word does not appear in this text” looks easy to check. It is not, because it depends entirely on what was searched. One of our pages claimed for two weeks that a treatment was not mentioned in a British guideline. It was, and it had a recommendation of its own. Since then, every claim of absence states which files were searched, with which search pattern, and on what date.
What we cannot guarantee
We read documents on a given date. A publisher can change a page without announcing it, and that happens. So we date our readings. When a link has not been reopened since the text was written, we say so rather than let you believe otherwise.
Our interactive pages do not write their text twice
A decision tree is displayed in two ways, as a drawing and as a list. The text, however, is written only once: the list is the source, the drawing is derived from it. That prevents a correction made on one side from leaving the other behind, a flaw that is never visible and is discovered months later. On every tree, a check compares the number of items written with the number of items drawn. The two must be equal.
When a source is closed, we say so instead of working around it
A source is sometimes closed at the moment we write: an article behind a paywall, supplementary material withdrawn, an agency site down. We do not fill the gap with a secondary source. We write that the item cannot be verified against the primary source, with the date, and that statement stays on the page.
Psychiatry Evidence Base, evidence-based psychiatry, explained with rigour.
