Evidence Review
Diagnostic Thresholds, Prevalence, and Overdiagnosis
What the evidence actually establishes
On this page23 sections
Companion reviews: Antipsychotics and the Evidence, ADHD — stimulants and the evidence, Depression — antidepressants and the evidence, Anxiety — a critical review of the evidence, and Pathologising the Disposition.
Executive summary
Two narratives dominate public discussion of psychiatric diagnosis: that rising diagnosis rates reflect a genuine epidemic driven by worsening child mental health (the "kids are sicker" story), and its inversion, that psychiatric diagnoses are social constructions with no biological reality (the "diagnoses are fake" story). Both are overclaims, and the evidence supports neither.
Eight things the evidence actually establishes:
-
ADHD prevalence is stable when standardized diagnostic procedures are used. Meta-regression of 135 studies found no association between year of study and prevalence estimates over three decades. Worldwide-pooled prevalence is 5.29% (Polanczyk et al., 2007) to 7.6% in children aged 3–12 years (Salari et al., 2023). Variability is explained by methodological factors: diagnostic criteria (DSM-III vs DSM-IV vs DSM-5), requirement for impairment, and source of information (parent only, teacher only, "or" rule, "and" rule, best-estimate procedure). Geographic location was not associated with variability when methodological factors were controlled (Polanczyk et al., 2014).
-
Diagnostic thresholds are rules, not natural boundaries. A DSM-5 ADHD diagnosis requires six symptoms (five for ages 17+), present before age 12, in two or more settings, with clear evidence of functional impairment. Change any element—symptom count, age-of-onset, impairment requirement, or informant rule—and prevalence changes, even if the underlying distribution of traits remains stable. Studies without impairment requirements report higher prevalence than those requiring impairment. Studies using "or" rules (parent or teacher endorsement) report higher prevalence than "and" rules or best-estimate procedures (Polanczyk et al., 2007).
-
The youngest children in a classroom are diagnosed with ADHD 60% more often than the oldest. Children born in the month before their state's kindergarten eligibility cutoff are diagnosed at 8.4%, versus 5.1% for those born in the month after. By fifth and eighth grades, the youngest children are nearly twice as likely as their oldest classmates to regularly use stimulants. Teachers' assessments of ADHD symptoms are strongly influenced by a child's age relative to classmates; parental assessments show weaker associations. This pattern suggests that many diagnoses are driven by developmental immaturity misinterpreted as pathology (Elder, 2010).
-
Overdiagnosis of ADHD is supported by evidence, particularly for milder cases. A systematic scoping review of 334 studies found substantial evidence of a reservoir of potentially diagnosable ADHD (104 studies), evidence that diagnoses had increased (45 studies), and that additional cases were milder (25 studies). Only 5 studies evaluated benefits and harms specifically in milder cases; these supported a hypothesis of diminishing returns where harms may outweigh benefits for youths with milder symptoms (Kazda et al., 2021).
-
DSM-5 field trial reliability varies dramatically by disorder. Intraclass kappa coefficients for test-retest reliability: ADHD 0.61 (very good), major depressive disorder 0.28 (questionable), generalized anxiety disorder 0.20 (questionable). Reliability categories: very good (0.60–0.79), good (0.40–0.59), questionable (0.20–0.39), unacceptable (<0.20). Among 23 diagnoses tested in adults and children, five were very good, nine good, six questionable, and three unacceptable (Regier et al., 2013).
-
What moves the prevalence needle is not underlying pathology but methodological choices. Thomas et al. (2015) found pooled prevalence of 7.2% across 175 studies (over 1 million children), but estimates ranged from 0.2% to 34%. Prevalence was lower in DSM-III-R studies than DSM-IV (p = .03) and lower in Europe than North America (p = .04). Few studies used population sampling with random selection; most sampled single towns or regions. The wide range reflects not geographic or temporal differences in true prevalence, but differences in how diagnostic criteria were operationalized.
-
Diagnostic expansion happens at the margins. DSM criteria have broadened over successive editions: DSM-IV (1994) required onset by age 7; DSM-5 (2013) moved this to age 12. Each relaxation captures cases previously excluded—not because the underlying condition has worsened, but because the boundary has moved. The Salari et al. (2023) finding that DSM-5 prevalence is higher than earlier criteria confirms this mechanism.
-
The two-informant problem is structural, not solvable. Parent-teacher agreement on ADHD ratings correlates at 0.34–0.64 (Strengths and Difficulties Questionnaire) and 0.17–0.60 (CBCL). Measurement invariance testing shows parents and teachers rate different behaviors in different settings, not the same behavior with noise. Combining informants with an "or" rule inflates prevalence; an "and" rule deflates it; a best-estimate procedure introduces clinician judgment as a variable. There is no methodologically neutral solution to this problem.
What the evidence does not establish:
- That psychiatric diagnoses lack validity or are "social constructions"
- That rising diagnostic and treatment rates reflect a true increase in underlying pathology when standardized procedures are used
- That children who meet diagnostic criteria do not experience genuine impairment
- Which diagnostic threshold (six symptoms vs five; "and" rule vs "or" rule; impairment required vs not) produces the "correct" prevalence
- That overdiagnosis is explained by pharmaceutical marketing alone (relative-age effects occur independently of industry influence)
The structural problem that conditions all of this: Psychiatric diagnoses are defined by symptom counts and clinician judgment, not by biomarkers or laboratory tests. The boundary between "disorder" and "no disorder" is set by committee consensus, not by natural discontinuities in the data. Diagnostic thresholds determine prevalence; prevalence does not validate thresholds. When the youngest-in-class are diagnosed 60% more often than the oldest, that is not measurement error—it is the system working as designed, comparing children within grades rather than across ages. The evidence establishes that diagnostic practice matters enormously, that thresholds are consequential choices rather than discoveries, and that prevalence estimates are as much about methodology as about underlying pathology.
How to read this review
The single most important heuristic
A prevalence figure without a named diagnostic method is not a finding. ADHD prevalence is 5.29% (Polanczyk 2007), 7.2% (Thomas 2015), or 7.6% (Salari 2023) depending on which studies are pooled, which DSM edition is used, whether impairment is required, and which informant rule is applied. Any prevalence estimate quoted without specifying diagnostic criteria, impairment requirements, and informant rules should be assumed to be the most favorable (or alarming) version available.
Three structural problems
Diagnostic thresholds are arbitrary. The choice of six symptoms versus five, age-of-onset by 7 versus 12, impairment required versus not, and "and" rule versus "or" rule are committee decisions, not empirical discoveries. Each choice moves the boundary and changes who gets diagnosed. There is no biomarker, no laboratory test, and no natural discontinuity that validates one threshold over another.
Relative standards generate diagnosis variance. When teachers compare a 5-year-old to 6-year-old classmates, developmental immaturity looks like inattention and impulsivity. The 60% diagnostic difference between youngest-in-class and oldest-in-class (Elder 2010) is not measurement error—it is the predictable result of asking teachers to rate children relative to peers rather than relative to age-normed standards.
Informant disagreement is not noise. Parents and teachers observe different behaviors in different settings (home versus classroom). Their ratings disagree not because one is wrong but because behavior is context-dependent. Diagnostic systems that require convergence across informants produce different prevalence estimates than systems that accept any informant's endorsement. Neither approach is methodologically neutral.
Conventions
Claims are marked [contested] where the literature genuinely disagrees and [unverified] where a figure is widely repeated but could not be traced to a primary source in preparing this review.
Confidence intervals are omitted where they could not be verified against the primary text, rather than reconstructed.
Part I — What diagnostic thresholds are, and why prevalence moves when thresholds move
1.1 The DSM-5 criteria for ADHD: a symptom-count model
The DSM-5 diagnostic criteria for ADHD require:
- Six or more symptoms of inattention and/or six or more symptoms of hyperactivity-impulsivity (five or more for individuals 17 years or older)
- Symptoms present before age 12 years
- Symptoms present in two or more settings (e.g., at home, school, or work)
- Clear evidence that symptoms interfere with or reduce the quality of social, academic, or occupational functioning
- Symptoms not better explained by another mental disorder
Each element is a threshold. Change the symptom count from six to five, and more children qualify. Move the age-of-onset from 7 (DSM-IV) to 12 (DSM-5), and more qualify. Require impairment, and fewer qualify than if impairment is optional. Require parent and teacher endorsement ("and" rule), and fewer qualify than if parent or teacher endorsement suffices ("or" rule).
These are not discoveries about the natural structure of ADHD. They are operationalization choices made by the DSM-5 Neurodevelopmental Disorders Work Group. Different choices produce different prevalence estimates, even when the underlying distribution of attention and activity traits in the population remains unchanged.
1.2 How threshold changes move prevalence: the DSM-IV to DSM-5 transition
DSM-IV (1994) required that "some hyperactive-impulsive or inattentive symptoms that caused impairment were present before age 7 years." DSM-5 (2013) changed this to "several inattentive or hyperactive-impulsive symptoms were present prior to age 12."
This change was not driven by new evidence that ADHD symptoms typically begin between ages 7 and 12. It was driven by field-trial data showing that the age-7 requirement excluded adults who met all other criteria but could not recall whether their symptoms began before age 7 or between ages 7 and 12. The Work Group decided that requiring recall of symptom onset before age 7 was unreliable and overly restrictive.
The result: children and adults with symptom onset between ages 7 and 12 now qualify for an ADHD diagnosis. Salari et al. (2023) found that ADHD prevalence under DSM-5 criteria is higher than under DSM-IV or DSM-III-R criteria. This is not evidence that ADHD has become more common. It is evidence that the diagnostic boundary has moved.
1.3 Impairment requirements: the hidden variable
Polanczyk et al. (2007) found that studies without a definition of impairment had significantly higher prevalence than those requiring impairment. This is one of the largest methodological effects in the meta-regression: presence or absence of an impairment criterion explained more variance in prevalence estimates than geographic location.
The DSM-5 criterion states: "Clear evidence that the symptoms interfere with, or reduce the quality of, social, academic, or occupational functioning." But "clear evidence" is not operationally defined. Some studies require clinician-rated impairment scales with threshold scores. Others require that the informant endorse "some problems" in at least one domain. Still others require no impairment assessment beyond the symptom count itself.
Studies that operationalize impairment rigorously report lower prevalence than studies that treat impairment as implicit in the symptom count. The prevalence difference is not measurement error. It reflects a substantive disagreement about whether meeting symptom count alone is sufficient for diagnosis, or whether evidence of functional impairment beyond symptom presence is required.
1.4 The informant problem: "or" rules, "and" rules, and best-estimate procedures
ADHD symptoms are present in "two or more settings"—typically home and school. Clinicians assess this by collecting parent and teacher ratings. But parent and teacher ratings often disagree. Correlations between parent and teacher ADHD symptom ratings range from 0.17 to 0.64 depending on the measure and sample (De Los Reyes et al., 2015, as cited in Gomez, 2019; see the ADHD review essay for extended discussion).
Diagnostic systems handle this disagreement in three ways:
- "Or" rule: Child meets criteria if parent or teacher endorses sufficient symptoms.
- "And" rule: Child meets criteria if parent and teacher both endorse sufficient symptoms.
- Best-estimate procedure: Clinician integrates parent, teacher, and child reports using clinical judgment.
Polanczyk et al. (2007) found that prevalence varies systematically by informant rule. Studies using "or" rules report higher prevalence than those using "and" rules or best-estimate procedures. Studies relying on parent-only or teacher-only information report prevalence estimates at the extremes of the distribution.
This is not a methodological flaw to be corrected. It is structural. Parent and teacher ratings disagree because children's behavior is context-dependent. Parents observe home behavior; teachers observe classroom behavior with 20–30 peers. A child may be inattentive and disruptive at school but not at home, or vice versa. Neither informant is "wrong"—they observe different behavior samples.
The choice of informant rule determines how many children with discordant ratings (high on one informant, low on the other) receive an ADHD diagnosis. There is no methodologically neutral answer to which rule is "correct."
Part II — ADHD prevalence: what the meta-regressions actually show
2.1 Polanczyk 2007: worldwide prevalence and the primacy of methodology
Polanczyk et al. (2007) conducted a systematic review and meta-regression of 102 studies (171,756 subjects) reporting ADHD/hyperkinetic disorder prevalence in subjects 18 years or younger according to DSM or ICD criteria.
The worldwide-pooled prevalence was 5.29% (95% CI 5.01–5.56).
But this figure conceals enormous variability. Prevalence estimates in individual studies ranged from less than 1% to over 10%. The meta-regression tested whether this variability was explained by geographic location, year of study, diagnostic criteria, impairment requirement, or source of information.
Key findings:
- Diagnostic criteria: Studies using DSM-III-R or ICD-10 criteria had significantly lower prevalence than those using DSM-IV criteria.
- Impairment requirement: Studies without a definition of impairment had significantly higher prevalence than those with an impairment requirement.
- Source of information: Studies using parent-only, teacher-only, or "or" rule had higher prevalence than those using best-estimate procedures. Studies using "and" rule had lower prevalence.
- Geographic location: After controlling for methodological factors, geographic location explained minimal variance. The only significant difference was between North America and both Africa and the Middle East. North America and Europe did not differ.
- Year of study: Not tested in the 2007 paper but tested in the 2014 update.
The authors concluded: "Our findings suggest that geographic location plays a limited role in the reasons for the large variability of ADHD/HD prevalence estimates worldwide. Instead, this variability seems to be explained primarily by the methodological characteristics of studies."
2.2 Polanczyk 2014: no evidence of increase over three decades
Polanczyk et al. (2014) updated the 2007 review, identifying 154 original studies and including 135 in multivariate analysis. The key question: Has ADHD prevalence increased over time?
The answer: No.
Meta-regression analyses tested the effect of year of study in the context of methodological variables (diagnostic criteria, impairment criterion, source of information) and geographic location.
Results:
- Methodological procedures investigated were significantly associated with heterogeneity of studies.
- Geographic location was not associated with variability in ADHD prevalence estimates.
- Year of study was not associated with variability in ADHD prevalence estimates.
The authors concluded: "Confirming previous findings, variability in ADHD prevalence estimates is mostly explained by methodological characteristics of the studies. In the past three decades, there has been no evidence to suggest an increase in the number of children in the community who meet criteria for ADHD when standardized diagnostic procedures are followed."
This finding refutes the "epidemic" narrative. Diagnostic and treatment rates have increased in many jurisdictions, but when researchers apply consistent diagnostic procedures across studies from different decades, prevalence estimates do not trend upward. The increase in diagnosed cases reflects changes in diagnostic practice (broader criteria, increased screening, reduced stigma), not an increase in the underlying condition.
2.3 Thomas 2015 and Salari 2023: prevalence depends on which studies you pool
Thomas et al. (2015) conducted a meta-analysis of 175 studies (179 prevalence estimates, over 1 million children) using DSM-III, DSM-III-R, or DSM-IV criteria.
Pooled prevalence: 7.2% (95% CI 6.7–7.8).
Why is this higher than Polanczyk 2007's 5.29%? Because the study pools different sets of studies with different methodological characteristics. Thomas et al. included studies published through 2013 and required that studies report point prevalence using DSM criteria. The resulting study pool had a different mix of diagnostic editions, impairment requirements, and informant rules than Polanczyk 2007.
Salari et al. (2023) conducted a meta-analysis of 53 studies of children aged 3–12 years.
Pooled prevalence: 7.6% (95% CI 6.1–9.4%).
Again, this figure is not inconsistent with Polanczyk 2007 or Thomas 2015. It reflects a different study pool with a narrower age range and a higher proportion of DSM-5 studies. Salari et al. explicitly note that "the prevalence of ADHD in children and adolescents according to the DSM-V criterion is also higher than previous diagnostic criteria."
The pattern that emerges: Prevalence estimates cluster around 5–8% depending on methodological choices. The variance between meta-analyses is explained by which studies are included, which DSM edition predominates, how impairment is defined, and which informant rules are used. Within any given methodological framework, prevalence is stable across time and geography. Change the framework, and prevalence changes—not because the underlying condition has become more or less common, but because the boundary has moved.
Part III — Overdiagnosis: what the scoping review actually found
3.1 Kazda 2021: a framework for detecting overdiagnosis in ADHD
Kazda et al. (2021) conducted a systematic scoping review of 334 studies using a five-question framework for detecting overdiagnosis in noncancer conditions:
- Is there potential for increased diagnosis? (reservoir)
- Has diagnosis actually increased?
- Are additional cases subclinical or low risk? (milder symptoms)
- Have some additional cases been treated?
- Might harms outweigh benefits of diagnosis (5a) and treatment (5b)?
Important: This review is not a prevalence meta-analysis. It does not report a "7.6% prevalence" figure. The 7.6% figure comes from Salari et al. (2023), a different paper. Kazda et al. reviewed 334 studies spanning 1979–2020 to evaluate evidence for overdiagnosis.
Key findings:
Question 1 (Reservoir): 104 studies provided substantial evidence of a reservoir of potentially diagnosable ADHD. Prevalence varies widely depending on diagnostic criteria, impairment requirements, and informant rules, indicating that a large number of children fall near the diagnostic boundary and could be diagnosed or not depending on how criteria are operationalized.
Question 2 (Diagnosis increase): 45 studies provided evidence that ADHD diagnosis rates have increased over time. This does not contradict Polanczyk 2014's finding of stable prevalence when standardized procedures are used. Diagnosis rates (the proportion of children receiving a clinical ADHD diagnosis) can increase even if true prevalence (the proportion meeting standardized diagnostic criteria) remains stable. The increase reflects changes in diagnostic practice, awareness, and access to services.
Question 3 (Milder cases): 25 studies showed that additional cases diagnosed in recent years are on the milder end of the ADHD spectrum. Children who would not have been diagnosed under older, more restrictive criteria are now diagnosed under broader criteria.
Question 4 (Treatment): 83 studies showed that pharmacological treatment of ADHD has increased over the same period.
Question 5 (Harms vs benefits in milder cases): This is the critical question for overdiagnosis. Only 5 studies evaluated benefits and harms specifically among youths with milder symptoms. These studies supported a hypothesis of diminishing returns: for youths with milder symptoms, the harms of diagnosis and treatment may outweigh the benefits.
The authors' conclusion: "This review found evidence of ADHD overdiagnosis and overtreatment in children and adolescents. Evidence gaps remain and future research is needed, in particular research on the long-term benefits and harms of diagnosing and treating ADHD in youths with milder symptoms."
3.2 What "overdiagnosis" means in this context
Kazda et al. define overdiagnosis as occurring when "a person is clinically diagnosed with a condition, but the net effect of the diagnosis is unfavorable." This is distinct from misdiagnosis (diagnosing ADHD when another condition is present) and false-positive diagnosis (diagnosing ADHD when symptoms are transient and resolve without intervention).
Overdiagnosis occurs when:
- A child meets diagnostic criteria (so the diagnosis is not technically "wrong"), and
- The child is on the mild end of the symptom spectrum, and
- The harms of diagnosis and treatment (labeling effects, medication side effects, opportunity costs) outweigh the benefits for that child.
The evidence supports overdiagnosis in this specific sense: there is a reservoir of children near the diagnostic boundary (Question 1), diagnosis rates have increased (Question 2), the additional cases are milder (Question 3), and the limited evidence on harms versus benefits in milder cases suggests diminishing or negative returns (Question 5).
Critically, this finding does not imply that ADHD is not a valid disorder, that children with severe symptoms do not benefit from diagnosis and treatment, or that the diagnostic criteria themselves are scientifically invalid. It implies that applying those criteria at population scale, with imperfect operationalization and informant disagreement, captures cases at the margins where net benefit is uncertain.
Part IV — Relative age effects: when developmental immaturity is mistaken for pathology
4.1 Elder 2010: the youngest in class are diagnosed 60% more often
Elder (2010) exploited state-level variation in kindergarten eligibility cutoff dates to test whether children's ADHD diagnosis rates depend on their age relative to classmates.
The natural experiment: A child born in October may begin kindergarten at age 5 if the state cutoff is December 1, but must wait until age 6 if the cutoff is September 1. Children born just before a cutoff become the youngest in their grade; those born just after become the oldest in the next grade.
Key findings:
- 8.4% of children born in the month before their state's cutoff (the youngest in their grade) were diagnosed with ADHD.
- 5.1% of children born in the month after the cutoff (the oldest in the next grade) were diagnosed with ADHD.
- This represents a 60% higher diagnosis rate among the youngest versus the oldest.
The discontinuity implies that the ADHD diagnosis rate among the youngest children in a classroom is 5.4 percentage points higher than it would have been if those children had waited an additional year to begin kindergarten. Given a baseline diagnosis rate of 6.4% in the sample, this is a substantial effect.
Teacher versus parent assessments:
- A child's birth date relative to the eligibility cutoff strongly influences teachers' assessments of whether the child exhibits ADHD symptoms.
- The same birth-date discontinuity is only weakly associated with parental assessments.
This pattern suggests that many diagnoses are driven by teachers' perceptions of poor behavior among the youngest children in a classroom. Teachers compare children within grades, not across ages. A 5-year-old in a class of 6-year-olds looks inattentive and impulsive relative to peers, even if the behavior is age-appropriate.
Long-term effects:
By fifth and eighth grades, the youngest children are nearly twice as likely as their oldest classmates to regularly use stimulants prescribed to treat ADHD. The initial diagnostic bias persists.
Elder notes that if these patterns are driven entirely by inappropriate diagnoses and treatment among the youngest children in a grade, the estimates imply that roughly 20% of the 2.5 million children using stimulants in 2010 were misdiagnosed.
4.2 Why relative-age effects are not measurement error
The standard interpretation of relative-age effects is that they reflect "inappropriate" diagnosis driven by teachers' misperception. But this framing obscures the structural problem: the diagnostic system itself asks teachers to compare children relative to same-grade peers.
DSM-5 criteria do not instruct teachers to rate whether a child's attention and activity are developmentally appropriate for the child's age. They instruct teachers to rate whether the child exhibits symptoms that "interfere with or reduce the quality of social, academic, or occupational functioning." In a classroom setting, "interfere with functioning" is inevitably judged relative to how other children in the class are functioning.
The youngest children in a grade are, on average, less mature than their older classmates. They have shorter attention spans, less impulse control, and more difficulty sitting still—not because they have a neurological disorder, but because they are younger. When teachers rate these behaviors on ADHD symptom checklists and compare ratings across children in the same classroom, the youngest children score higher.
This is not a flaw in teacher ratings. Teachers are doing exactly what the diagnostic system asks them to do: rate behavior and impairment relative to typical functioning in the child's environment. The flaw is in the system's failure to account for age-for-grade when interpreting those ratings.
4.3 Implications for prevalence estimates
Relative-age effects demonstrate that ADHD diagnosis rates are sensitive to subjective, within-group comparisons. Prevalence estimates that rely on teacher ratings in school-based samples will systematically inflate diagnosis rates among the youngest children in each grade.
Elder's estimates suggest that 20% of stimulant prescriptions in 2010 may have been driven by relative-age effects. If we apply this estimate to the 5.29% worldwide pooled prevalence from Polanczyk 2007, it implies that roughly 1 percentage point of the 5.29% (approximately 20%) may be attributable to developmental immaturity rather than ADHD.
This does not mean the remaining 80% of diagnoses are "correct" in some absolute sense. It means that even after controlling for diagnostic criteria, impairment requirements, and informant rules, an additional source of variance—relative age within grade—systematically biases prevalence estimates upward for the youngest children.
Part V — DSM-5 field trials: how reliably can clinicians agree on who has a disorder?
5.1 The DSM-5 field trial design
Regier et al. (2013) conducted test-retest reliability trials for DSM-5 diagnoses across 11 sites in the United States and Canada. The design: two clinicians independently interview the same patient on separate occasions and assign DSM-5 diagnoses. The intraclass kappa coefficient measures agreement.
Reliability categories:
- Very good: κ = 0.60–0.79
- Good: κ = 0.40–0.59
- Questionable: κ = 0.20–0.39
- Unacceptable: κ < 0.20
A total of 2,246 patients were enrolled; over 86% completed two diagnostic interviews. Clinicians were trained using methods comparable to what would be available after DSM-5 publication (i.e., not research-level training).
5.2 Results: reliability varies dramatically by disorder
Among 23 adult and child/adolescent diagnoses with adequate sample sizes:
Very good (κ = 0.60–0.79):
- Post-traumatic stress disorder (0.67)
- Major neurocognitive disorder (0.78)
- Complex somatic disorder (0.61)
- Autism spectrum disorder (0.69, child/adolescent)
- ADHD (0.61, child/adolescent)
Good (κ = 0.40–0.59):
- Schizophrenia (0.46)
- Bipolar I disorder (0.54)
- Alcohol use disorder (0.40)
- Borderline personality disorder (0.54)
- Avoidant/restrictive food intake disorder (0.48, child/adolescent)
- Oppositional defiant disorder (0.40, child/adolescent)
Questionable (κ = 0.20–0.39):
- Major depressive disorder (0.28, adult and child/adolescent)
- Generalized anxiety disorder (0.20)
- Antisocial personality disorder (0.21)
- Mild neurocognitive disorder (0.28)
- Disruptive mood dysregulation disorder (0.25, child/adolescent)
Unacceptable (κ < 0.20):
- Mixed anxiety-depressive disorder (0.00, adult; 0.05, child/adolescent)
- Nonsuicidal self-injury (<0.20, child/adolescent)
5.3 What low reliability means for prevalence
Major depressive disorder and generalized anxiety disorder—two of the most commonly diagnosed conditions—showed questionable reliability. Independent clinicians agree on these diagnoses at rates barely above chance.
This is not solely a clinician-training problem. It reflects genuine ambiguity in where the diagnostic boundary lies. DSM-5 criteria for major depressive disorder require five of nine symptoms for at least two weeks, with at least one symptom being depressed mood or loss of interest. But symptoms like "fatigue," "difficulty concentrating," and "feelings of worthlessness" are continuous dimensions, not binary present/absent states. Different clinicians will threshold these symptoms differently.
Generalized anxiety disorder (κ = 0.20) performed even worse. The diagnostic boundary between "excessive" worry and "normal" worry is inherently subjective, and the requirement that worry be "difficult to control" adds a second layer of subjectivity.
By contrast, ADHD (κ = 0.61) showed very good reliability in the child/adolescent sample. Why? ADHD criteria are more behavioral and observable (e.g., "often fidgets," "often leaves seat") than the criteria for depression or anxiety. Teachers and parents can more reliably endorse whether a child "often interrupts others" than whether an adult's worry is "excessive" or "difficult to control."
But even for ADHD, κ = 0.61 is in the "very good" range—not perfect agreement. And this is test-retest reliability within a single site using DSM-5 criteria. Cross-site reliability, or reliability across different diagnostic criteria (DSM-IV vs DSM-5), would be lower.
5.4 Implications for epidemiological prevalence estimates
Field-trial reliability sets an upper bound on the precision of prevalence estimates. When two trained clinicians independently interviewing the same patient produce chance-corrected agreement of κ = 0.61 for ADHD (very good) or κ = 0.28 for depression (questionable), population prevalence estimates derived from those diagnoses will reflect that measurement uncertainty.
Prevalence studies typically use structured diagnostic interviews or rating scales administered by trained research assistants, not independent clinical re-interviews. The reliability of these instruments is often higher than field-trial reliability, because the instruments standardize the diagnostic process. But the trade-off is that structured instruments may lack clinical validity—they capture symptom counts reliably but miss the clinical context that determines whether symptoms constitute a disorder.
The field trials establish that diagnostic reliability is a limiting factor in psychiatric epidemiology. Even for well-defined, behaviorally anchored disorders like ADHD, chance-corrected agreement is κ = 0.61 (very good)—not perfect. For mood and anxiety disorders, chance-corrected agreement is in the questionable range (κ = 0.28 for MDD, κ = 0.20 for GAD). Prevalence estimates derived from these diagnoses will inherit that measurement uncertainty.
Part VI — The two overclaims, and what replaces them
6.1 The "epidemic" narrative: "Kids are sicker now"
The claim: Rising ADHD diagnosis and treatment rates reflect a true increase in the underlying condition. Children today are more distracted, more impulsive, and more hyperactive than children in previous generations, driven by screens, social media, environmental toxins, or other modern stressors.
What the evidence shows: Polanczyk et al. (2014) found no association between year of study and ADHD prevalence estimates in a meta-regression of 135 studies over three decades. When standardized diagnostic procedures are used, prevalence is stable. The increase in diagnosis and treatment rates reflects changes in diagnostic practice (broader criteria, increased awareness, reduced stigma, more screening), not an increase in the underlying condition.
What survives: Diagnosis and treatment rates have increased, but this does not imply that more children meet standardized diagnostic criteria. It implies that more children who meet criteria are being identified and treated, and that diagnostic criteria have broadened (DSM-IV to DSM-5), capturing cases previously excluded.
6.2 The invalidation narrative: "Diagnoses are fake"
The claim: ADHD is not a real disorder but a medicalization of normal childhood behavior. Psychiatric diagnoses are social constructions with no biological reality, driven by pharmaceutical marketing and societal intolerance of active, non-conforming children.
What the evidence shows:
- ADHD symptoms cluster consistently across cultures and diagnostic systems (Polanczyk 2007: geographic location explains minimal variance after controlling for methodology).
- Children who meet diagnostic criteria show functional impairment in academic, social, and family domains (Kazda 2021: studies consistently report impairment in diagnosed cases, though impairment is less clear in milder cases).
- Diagnostic reliability for ADHD is κ = 0.61, in the "very good" range—better than most psychiatric diagnoses (Regier 2013).
- The prevalence range (5.3–7.6%) is narrow relative to the methodological variance that could inflate it. If ADHD were purely a social construction, we would expect prevalence to vary wildly by culture, not cluster around 5–8% after controlling for methodology.
What survives: ADHD symptoms are dimensionally distributed. The boundary between "disorder" and "no disorder" is set by symptom-count thresholds and impairment requirements, not by natural discontinuities. Different threshold choices produce different prevalence estimates, and those choices are made by diagnostic committees, not discovered empirically. But this does not mean the symptoms are not real or that they do not cause impairment.
6.3 The version that survives
Diagnostic thresholds are consequential choices, not empirical discoveries. The DSM-5 decision to require six symptoms (not five or seven), age-of-onset by 12 (not 7 or 15), and impairment (variously operationalized) determines who gets diagnosed. Change any element, and prevalence changes, even if the underlying trait distribution is stable. These are committee decisions informed by field-trial data, expert consensus, and clinical experience—but they are not validated by biomarkers or natural boundaries in the data.
Prevalence estimates are as much about methodology as about pathology. The range from 5.3% to 7.6% across meta-analyses reflects different mixes of studies with different diagnostic criteria, impairment requirements, and informant rules. Within any methodological framework, prevalence is stable across time and geography. Change the framework, and prevalence changes. Prevalence does not validate thresholds; thresholds determine prevalence.
Relative-age effects demonstrate that diagnosis rates are sensitive to subjective comparisons. The youngest children in a classroom are diagnosed 60% more often than the oldest, not because of neurological differences, but because teachers compare them to older peers (Elder 2010). This is not measurement error—it is the system working as designed, comparing children within grades rather than adjusting for age. The diagnostic system does not account for this source of variance, so it propagates into prevalence estimates.
Overdiagnosis occurs at the margins, where harms may outweigh benefits. Kazda et al. (2021) found evidence of a diagnostic reservoir, increasing diagnosis rates, milder additional cases, and limited evidence on net benefit for those milder cases. Overdiagnosis in this sense does not invalidate the diagnostic category. It suggests that applying categorical diagnostic criteria to dimensional traits, without biomarkers or functional outcome data, captures cases where intervention may not improve (or may worsen) outcomes.
The two folk models—"epidemic" and "fake diagnoses"—are both wrong. Kids are not "sicker now" (prevalence is stable when methods are constant). Diagnoses are not "fake" (symptoms cluster reliably; impairment is real; diagnostic reliability is acceptable). What has changed is diagnostic practice: broader criteria, more screening, increased treatment access. What remains unchanged is that diagnostic thresholds are normative choices, not empirical discoveries, and that prevalence moves when thresholds move.
References
Primary sources cited
Elder, T. E. (2010). "The importance of relative standards in ADHD diagnoses: Evidence based on exact birth dates." Journal of Health Economics, 29(5), 641–656.
Kazda, L. et al. (2021). "Overdiagnosis of Attention-Deficit/Hyperactivity Disorder in Children and Adolescents: A Systematic Scoping Review." JAMA Network Open, 4(4), e215335.
Polanczyk, G. et al. (2007). "The Worldwide Prevalence of ADHD: A Systematic Review and Metaregression Analysis." American Journal of Psychiatry, 164(6), 942–948.
Polanczyk, G. V. et al. (2014). "ADHD prevalence estimates across three decades: an updated systematic review and meta-regression analysis." International Journal of Epidemiology, 43(2), 434–442.
Regier, D. A. et al. (2013). "DSM-5 Field Trials in the United States and Canada, Part II: Test-Retest Reliability of Selected Categorical Diagnoses." American Journal of Psychiatry, 170(1), 59–70.
Salari, N. et al. (2023). "The global prevalence of ADHD in children and adolescents: a systematic review and meta-analysis." Italian Journal of Pediatrics, 49, 48.
Thomas, R. et al. (2015). "Prevalence of attention-deficit/hyperactivity disorder: a systematic review and meta-analysis." Pediatrics, 135(4), e994–1001.
Secondary sources cited
De Los Reyes, A. et al. (2015). Cited in Gomez (2019) for parent-teacher agreement correlations on ADHD rating scales.
Gomez, R. (2019). "Is the endorsement of the Attention Deficit Hyperactivity Disorder symptom criteria ratings influenced by informant assessment, gender, age, and co-occurring disorders? A measurement invariance study." International Journal of Methods in Psychiatric Research, 28(3), e1794. (Cited in the companion ADHD essay for measurement invariance findings; correlations reported here are from the ADHD essay, which cites De Los Reyes 2015 via Gomez 2019.)
Companion essays
This essay provides the empirical grounding for Pathologising the Disposition, which examines the conceptual structure of deficit-seeking assessment frameworks. Together, they establish that diagnostic thresholds are normative choices applied to dimensional traits, and that prevalence estimates depend as much on those choices as on the distribution of underlying pathology.
For evidence reviews of specific treatments and outcomes, see:
About the author
Paul Stephen
Founder, Apatheia Labs
Evidence-governed research publication — Prosoche applied in the open.
Read next
More in Evidence- EssayAugust 2026
Adult ADHD and the Evidence
Prospective cohort studies find that childhood ADHD persistence to adulthood is low across named cohorts (Dunedin 5%, Pelotas 17.2%, E-Risk 21.9%), most adult-ADHD cases lack prospectively diagnosed childhood ADHD (Dunedin 90%, Pelotas 84.6%, E-Risk 67.5%), and after excluding substance use and other disorders, fewer than 1% of originally undiagnosed children meet late-onset criteria.
- EssayAugust 2026
ADHD Medications and the Evidence
Guanfacine and clonidine show child clinician SMDs −0.67 and −0.71 (Cortese 2018), but atomoxetine and viloxazine carry boxed warnings for suicidal ideation. Adult lisdexamfetamine Castells 2018 SMD −1.06; bupropion Verbeeck 2017 pooled −0.50 but clinician-only ns. Adult ER-MPH Boesen 2022 investigator SMD −0.42; IR-MPH Cândido 2021 investigator MD −20.70 but participant ns. Adult GXR Iwanami 2020 Japan diff −4.28; child GXR Sallee 2009 LS-mean −5.41 to −7.88.
- EssayAugust 2026
Antipsychotics and the Evidence
All antipsychotics beat placebo in acute trials (SMD 0.33–0.88), but 74% of CATIE participants discontinued within 18 months, dose reduction in first-episode patients doubled 7-year recovery rates versus maintenance treatment, and long-term observational data show better outcomes off medication—confounded by selection, but challenging the 'indefinite treatment for all' assumption.