Published on 14 September 2026
Complex PTSD: does the ITQ measure what ICD-11 describes?
In brief
The International Trauma Questionnaire is the instrument built to operationalise the two ICD-11 diagnoses, post-traumatic stress disorder and its complex form. This meta-analytic confirmatory factor analysis brings together 57 studies and 43,066 participants in 31 countries, and compares nine factor models. The model that fits the data best is not the one reproducing the ICD-11 structure, but a seven correlated first-order factor model in which affective dysregulation splits into two distinct dimensions, an affective hyperactivation component and an affective hypoactivation component. The ICD-11 model nonetheless keeps a good fit, and the authors conclude that the instrument remains adequate for assessing both diagnoses. A second finding bears immediately on the consultation: the reliability of several subscales of the short version is insufficient, that of affective dysregulation above all, while the two higher-order scores stay reliable. The data and the code are deposited in open access, the protocol was registered in advance, and only 38% of the authors approached supplied their data matrices.
The context
Since ICD-11 separated post-traumatic stress disorder from its complex form, the distinction has rested on three additional symptom clusters grouped under the name disturbances in self-organization: affective dysregulation, negative self-concept, and disturbances in relationships. The instrument meant to capture these dimensions established itself very quickly, in research as in the clinic. Its finalised version, released in 2018, carries twelve symptom items, two per cluster, plus six functional impairment items, eighteen items in all, each rated from 0 to 4. It has been translated into twenty-five languages. That brevity is a deliberate choice by its designers, who put clinical utility first by keeping a minimal number of indicators per dimension.
The question asked here is not whether the complex form exists. It is whether the instrument used to identify it behaves statistically the way the classification assumes. This is a psychometric question, and it has very concrete consequences for what a clinician is entitled to conclude from a score. The authors position their work by stating that no quantitative synthesis had so far examined the latent structure of this instrument, a claim restricted to that single object, and one that leaves aside the earlier qualitative reviews whose results they go on to discuss at length.
The study at a glance
| Population | |
| Participants who completed the International Trauma Questionnaire, all types of traumatic exposure, 31 countries, 43,066 people across 57 studies published between 2016 and 2024, and 58 distinct samples. Median sample size 307 (mean 743, range 44 to 4,944), mean age of the studies from 19.09 to 67.08 years (median 37.4), proportion of women from 6.20% to 100% (mean 62%) | |
| Intervention, in the sense of the analysis performed | |
| Meta-analytic confirmatory factor analysis by two-stage structural equation modelling, random effects for the main analysis, testing nine competing factor models. The subgroup analyses had to fall back on fixed-effects models, for want of a sufficient number of studies | |
| Comparator | |
| The ICD-11 model, that is two second-order factors, post-traumatic stress disorder and disturbances in self-organization, sitting above six symptom clusters, set in particular against a seven correlated first-order factor model separating two components of affective dysregulation, and against a six correlated first-order factor model | |
| Outcome | |
| Comparative goodness of fit across models on six indices and the chi-square, then score reliability through coefficient omega, and a search for moderators through subgroup analyses and meta-regression of factor loadings | |
| Design | |
| Systematic review following PRISMA, registered with PROSPERO under number CRD42023434942, eight databases searched in April 2023 and again in April 2024, no language restriction, quality appraised with a modified version of the QualSyst tool, data and code deposited in open access on OSF |
Quality control
| Point checked | Judgement |
|---|---|
| Pre-registration | Done |
| FindingProtocol lodged with PROSPERO under number CRD42023434942, which limits after-the-fact selection of the models reported. The PRISMA checklist sits in the supplementary material | |
| Size of the pooled sample | Solid |
| Finding43,066 participants, 57 studies, 31 countries, eight databases. Rarely reached in psychometrics | |
| Inter-rater agreement | High |
| FindingCohen’s kappa 0.91 for title and abstract screening, 0.93 for full texts, 0.89 for the quality appraisal. Disagreements settled by discussion, extraction discrepancies on 8 of 472 data points | |
| Quality of the included studies | Mixed |
| FindingModified QualSyst tool, four criteria. 79% of the studies (45 of 57) are rated high quality, 21% moderate quality, none low quality. But the data quality criterion is the weak point of the corpus: 27 studies (48.2%) are rated weak on it, mainly for failure to report missing data (24 studies) | |
| Heterogeneity | High |
| FindingAt the first stage of the analysis, Q = 18,230.74 on 3,586 degrees of freedom, p < 0.001, and I² between 11.15% and 97.49% depending on the correlation. Exact homogeneity is therefore rejected, which calls for caution in generalising | |
| Data transparency | Open |
| FindingData and scripts deposited in open access on OSF, analyses run in R with the metaSEM, OpenMx and lavaan packages, versions stated. A third party can reproduce the analysis | |
| Access to the raw data | Partial |
| FindingNo included publication provided a complete covariance matrix. 126 authors were approached about 185 studies, with 48 replies, that is 38%. An availability bias cannot be ruled out, since the teams that answer are not necessarily representative | |
| Funding and declared interests | Declared |
| FindingAcademic work, funded in part by an Australian Government Research Training Program Stipend and by the Dutch Research Council (project VI.Vidi.201.009). The authors declare no potential conflicts of interest with respect to the research, authorship and publication of this article | |
The findings
| Result | Value |
|---|---|
| Best-fitting model | Seven correlated factors |
| PEB readingChi-square 215.16 on 33 degrees of freedom, RMSEA 0.011, SRMR 0.012, TLI 0.993, CFI 0.997, BIC negative at −136.97, the lowest of the nine models. Every factor loading is substantial, from 0.75 to 0.92, and the correlations between factors run from 0.42 to 0.77 | |
| The ICD-11 model | Good fit, ranked third |
| PEB readingThe two-factor second-order model reaches RMSEA 0.020, SRMR 0.029, TLI 0.980, CFI 0.986, values that correspond to a good fit under the thresholds adopted. It comes third, behind the seven-factor model and the six correlated factor model. The correlation between the core PTSD factor and the disturbances in self-organization factor is 0.68, which makes them related but distinguishable. The authors write of this model that it “demonstrated a good fit, and for this reason, it should not be dismissed” | |
| What the seventh dimension is | A split in affective dysregulation |
| PEB readingAn affective hyperactivation component and an affective hypoactivation component behave statistically as two distinct dimensions rather than one. The superiority of this model replicates on the 22-item long form and on three modified versions of the instrument, which makes an artefact of item redundancy unlikely | |
| Reliability of the higher-order scores | Good to excellent |
| PEB readingIn the ICD-11 model, second-order omega reaches 0.84 for the core PTSD score and 0.87 for the disturbances in self-organization score on the 12-item version, 0.89 and 0.96 on the 22-item version. The thresholds the authors adopt are 0.70 to 0.79 for acceptable, 0.80 to 0.89 for good, 0.90 and above for excellent | |
| Reliability of the subscales | Insufficient for four of six |
| PEB readingOn the 12-item version, omega is 0.61 for re-experiencing, 0.69 for avoidance, 0.66 for sense of threat and 0.43 for affective dysregulation, all below the acceptable threshold. Only negative self-concept (0.84) and disturbances in relationships (0.71) reach it. This is the finding most directly usable in the consultation: a subscale score taken on its own does not carry information stable enough to act on | |
| The 22-item long form | Better, and still insufficient |
| PEB readingOmega values are higher there (0.77 for avoidance, 0.74 for sense of threat, 0.78 for disturbances in relationships), but affective dysregulation stays at 0.55, and its two components, once separated, at 0.49 for hyperactivation and 0.66 for hypoactivation, despite five and four items. Item count therefore does not explain everything | |
| Search for moderators | Significant effects, negligible magnitude |
| PEB readingThe meta-regression of factor loadings shows statistically significant moderation by the PTSD status of the sample (chi-square 61.80 on 12 degrees of freedom, p < 0.001, mean change −0.10), by crossing the probable complex PTSD threshold (chi-square 585.73, p < 0.001, mean change −0.11), by symptom severity and by the language of administration (mean changes −0.01). The authors adopt a threshold of 0.10 in absolute value below which a difference in loading is judged insubstantial, and conclude that these effects, significant as they are, do not alter the structure. Trauma type was not tested as a moderator | |
| Reliability by population | Lower in clinical samples |
| PEB readingPTSD samples show lower reliability estimates, from −0.06 to −0.15 depending on the subscale, and the groups reaching the probable complex PTSD threshold fall as far as −0.22 on re-experiencing. In other words, the instrument is least reliable at subscale level exactly where it is used most | |
Critical appraisal
| Domain | Judgement |
|---|---|
| Search and study selection | Low risk |
| FindingEight databases, a registered protocol, explicit criteria, no language restriction, double screening with high agreement | |
| Availability of the primary data | Moderate risk |
| FindingThe 38% response rate leaves open the possibility that the samples analysed differ systematically from those that could not be. The authors also acknowledge that the grey literature was not exhaustively covered | |
| Model comparison | Low risk |
| FindingNine models compared on six fit indices, with three sensitivity analyses, one excluding low and moderate quality studies, one excluding translations that were not back-translated, the third excluding studies rated poorly on missing data. The comparison is not restricted to the authors’ preferred model, and the ranking holds in all three cases | |
| Scope of the factorial result | Not to be overstated |
| FindingA better statistical fit does not demonstrate that a diagnostic category is wrong. The authors conclude the opposite: that the study affirms the adequacy of the instrument for assessing the two ICD-11 diagnoses, that the model matching that classification keeps a good fit, and that using its second-order factors may remain advantageous in clinical work until the conceptualisation of the complex form is refined. That qualification has to be preserved | |
| Measurement invariance | Not tested |
| FindingThe authors state explicitly that the focus of the paper was neither latent factor variances nor invariance testing, and that working from correlation matrices rather than covariance matrices did not allow it. Thirty-one countries represented therefore do not amount to a demonstration of measurement equivalence across languages and settings. The meta-regression by language of administration does not stand in for that test | |
| Subgroup analyses | Moderate risk |
| FindingThe PTSD subgroup holds only 9 studies and 6,151 participants and had to be analysed under fixed effects, which restricts generalisation beyond the included studies. The method used allows neither continuous moderators nor the simultaneous inclusion of several moderators, which exposes the analysis to confounding | |
Level of evidence
PEB appraisal: a strong method, and no Oxford level that fits. The Oxford scale has no heading for the structural validity of an instrument: its levels answer questions of treatment, diagnosis, prognosis or risk. No level is recorded here, and we prefer to say so rather than transpose a grading outside its domain. Confidence is high on the finding of insufficient reliability for the short subscales, which rests on a very large sample, a transparent method and open data. It is high as well on the better fit of the seven-factor model, with two reservations: this concerns the structure of the instrument and not the validity of the diagnostic category, and heterogeneity between studies is substantial. It is lower on transposing the result to one particular language setting, for want of an invariance analysis in this work.
The colleague test
What an experienced colleague would say if you put this study to them in two minutes, between two consultations.
“ I keep using the questionnaire, but I read the two higher-order scores and I no longer comment on a single subscale. The rest comes from the interview. ”
What this means in practice: a high score on one dimension of disturbances in self-organization is not enough to ground clinical reasoning. The two higher-order scores are.
What you can do with this
- Use the two higher-order scores of the instrument, the one for post-traumatic stress disorder and the one for disturbances in self-organization, and stay with them for any decision. Their reliability is good, that of the subscales is not.
- Do not base a conclusion on a subscale taken on its own in the short version, above all affective dysregulation, whose omega coefficient falls to 0.43. The sum of two items does not measure a dimension in a stable way.
- Remember that subscale reliability drops further in the most symptomatic populations, precisely those seen in specialist practice.
- If the aim is a detailed profile dimension by dimension, consider the 22-item long form rather than the short version, as the authors suggest, without expecting it to settle everything: affective dysregulation stays below the acceptable threshold there too.
- Explore affective dysregulation in both directions during the interview, what overflows and what shuts down, rather than as a single dimension.
- Always complete the self-report questionnaire with a clinical interview before settling on a diagnosis of the complex form.
- Be ready to answer a patient who has filled in the questionnaire: the tool points a direction, it does not make the diagnosis.
Frequently asked questions
Does this study call complex PTSD into question?
No. It examines the factor structure of the instrument used to measure it. The authors conclude that the instrument is adequate for assessing the two ICD-11 diagnoses, and they write of the model matching that classification that it “demonstrated a good fit, and for this reason, it should not be dismissed”.
Should the questionnaire be abandoned?
Nothing in this analysis suggests it. The authors describe it as a strong and brief measure, suitable for assessing the two ICD-11 diagnoses at the factor level. What is advised against is using the subscales in isolation, not using the tool.
What does insufficient reliability mean in practice?
That a subscale score varies too much, between items meant to measure the same thing, to carry a clinical decision on its own. With two items per dimension in the short version, one atypical answer moves the whole score.
Does the result hold for a translated version of the questionnaire?
Of the 57 included studies, most used a translated version, and the language of administration was tested as a moderator of factor loadings, with a negligible mean change. But measurement invariance was not tested in this work, and the authors say so themselves. The validation of any given language version has to be checked separately.
Will the seven-factor model replace ICD-11?
That is not what the study says. The authors suggest that future iterations of the ICD formulation of the complex form should specify the affective hyperactivation and hypoactivation dimensions. It is a reasoned proposal, not a decision about nomenclature.
Annotated bibliography
Kindred R, Jak S, Hamer R, Nedeljkovic M, Bates GW (2026). Evaluating the ICD-11 PTSD and Complex PTSD Constructs: A Meta-Analytic Confirmatory Factor Analysis of the International Trauma Questionnaire. Assessment, 33(4), 510-532. DOI 10.1177/10731911251340837 · PMID 40448311. Source study analysed here. Meta-analytic confirmatory factor analysis by two-stage structural equation modelling, 57 studies, 43,066 participants, 31 countries, protocol registered with PROSPERO under number CRD42023434942, data and scripts deposited on OSF. June 2026 issue, electronic publication 30 May 2025.
What was consulted. References checked on 13 August 2026 against the published version. The supplementary material, a single file gathering supplemental tables 1 to 22 and supplemental figures 1 and 2, was not consulted: it did not accompany the version available to us. No figure in this analysis comes from it. This article underwent an independent double reading.
