Reliability of psychiatric diagnoses in the 21st century
Mateo Boberg, Mads Gram Henriksen, … Léon Franzen, Stefan Borgwardt, …
DOI:
10.3389/fpsyt.2026.1891625
Abstract
Background:
Reliable diagnosis is a cornerstone of medical practice, yet concerns about diagnostic reliability have long haunted psychiatry. In the 1970s, the US–UK Diagnostic Project revealed striking international discrepancies in diagnostic practice. Although operationalized diagnostic criteria were introduced to improve reliability, contemporary evidence indicates that substantial problems persist. To investigate this, we conducted a large-scale international study of diagnostic reliability among medical doctors using standardized clinical case vignettes.
Material and methods:
In this cross−sectional study, medical doctors working in adult psychiatry across 19 countries from two continents assigned ICD−10 diagnoses to written cases. Each doctor was randomly assigned two of nine available cases. Inter−rater reliability among medical doctors was assessed using Krippendorff’s α. Diagnostic variability at case-level was quantified using Shannon entropy. We also calculated diagnostic accuracy as concordance with predefined best-estimate reference diagnoses.
Results:
A total of 1,038 medical doctors provided 1,902 diagnostic assessments across nine cases. Inter−rater reliability among medical doctors was modest (Krippendorff’s α = 0.48), indicating that diagnostic agreement between two doctors occurred in roughly 55% of cases. Overall diagnostic accuracy was 66%, with the lowest accuracy observed for cases depicting schizophrenia spectrum disorders. Diagnostic disagreements followed systematic patterns, with schizophrenia cases frequently misclassified as OCD, personality disorder, or bipolar disorder.
Conclusions:
Diagnostic inter−rater reliability among medical doctors remains limited, even under standardized conditions. The observed patterns of misclassification suggest that the variability in diagnostic assessments reflects systematic differences in clinical interpretation rather than random error, pointing to challenges in assessing psychopathology as well as ambiguities in psychiatric classification.