A critical review of the evidence — and of forty years spent looking for a pathophysiology that has not turned up.
Compiled 27 July 2026 · 115 primary sources · 17 load-bearing claims adversarially verified · companion to Stress, Energy and the Capacity to Function
The short version
The headline findings are mostly negative — and that is the finding
The stress review found a field where the folk model was wrong but a better mechanistic story was waiting underneath. Anxiety is harder. Here the folk model is also wrong, and what replaces it is largely a set of well-powered nulls, demonstrated confounds and outright reversals. Two treatments work. Almost nothing else survives contact with an honest comparator.
0
Dot-probe variants, out of 36 tested in 9,600 people, with internal reliability above zero — the measure two decades of attentional-bias theory was built on
Xu et al. 2025, Clin Psychol Sci
rg = 0.91
Genetic correlation between anxiety and depression in the largest GWAS to date (122,341 cases)
Strom et al. 2026, Nature Genetics
−0.56
Individual CBT for social anxiety against a psychological placebo — versus −1.19 against a waitlist. Roughly half the advertised effect
Mayo-Wilson et al. 2014, Lancet Psychiatry
p = 0.28
Cortisol stress reactivity in anxiety. The HPA axis is null — and where it has a sign, it points down
Zorn et al. 2017
45–65%
Response rate to first-line treatment. A third to a half of people do not respond at all
Bandelow et al. 2014, 403 RCTs
40 mice
The entire empirical basis of Yerkes–Dodson — with two per data point in the only condition that produced the famous curve
Yerkes & Dodson 1908, original text
The structural conclusion that runs underneath all of it. Quantitative structural models place GAD closer to major depression than to panic or the phobias. Intolerance of uncertainty is transdiagnostic and dissolves into neuroticism. The general psychopathology factor performs no better than a raw count of comorbid diagnoses. It is at least as likely that the category is wrong as that the instruments are — and you will not find the biology of a thing that is not a kind.
The single most useful heuristic
An effect size without a comparator is not a finding
Social anxiety disorder, 101 trials, 13,164 participants. This is the cleanest demonstration in psychiatry of how much comparator choice does. Every intervention was measured against a waitlist; only two were also measured against a credible placebo.
Effect sizes before and after methodological correction
Standardised mean difference · larger = better · tap or hover to see what the correction was
as first reported after correctioncorrections: honest comparator · high-quality trials only · publication-bias adjustment
↔ scroll the chart sideways, or use the table view
Every arrow points the same way, and most of them halve. The corrections are of three kinds — swapping a waitlist for a credible placebo, restricting to high-quality trials, and adjusting for publication bias — and they are not exotic; each is standard practice. The two rows that matter most are the top two: individual CBT and SSRIs/SNRIs were the only intervention classes in a 101-trial network to beat their appropriate placebo, and both lose about half their advertised magnitude in doing so. At the bottom, three interventions cross into non-significance and one lands essentially at zero.
The working rule. Any effect size quoted in this field without its comparator should be assumed to be the waitlist number, and mentally halved. That single habit will correct more errors than any other piece of methodological knowledge here.
The pattern
Four domains, one arc
A striking small-sample result, a mechanistic theory built on it, sometimes an intervention built on the theory — and then one of three endings. This recurs so consistently that it is worth treating as the field's central finding rather than as a series of accidents.
Ending 1 — a well-powered null
Interoceptive accuracy: no association, 55 studies, pre-registered
Cortisol reactivity: −0.101, p = 0.28
ENIGMA GAD: "no effect of GAD on brain structure"
GABA on MRS: g = −0.26, p = 0.52
Volatility learning: 8 experiments, N = 820, no systematic effect
CRP: association reverses after adjusting for depression
HF-HRV: the respiration confound is differential — present in patients, absent in controls
Intolerance of uncertainty: no unique association after neuroticism facets
Days out of role: crude figure is ~3× the adjusted — comorbidity carries it
Attainment: HR 1.56 → 0.96 in discordant twins
Ending 3 — a direct reversal
Serotonin is elevated, not depleted, in social anxiety
Oestradiol impaired extinction recall in the best-powered RCT
CRP is protective on Mendelian randomisation (OR 0.87)
Amygdala and BNST respond to the opposite threat types to the model
Worry raises autonomic arousal (k = 60, every sign)
Fluoxetine increased anxiety-like behaviour in the standard rodent assay
The most telling single fact in the mechanisms literature. Across 814 studies, only 2 of 17 commonly used elevated-plus-maze measures reliably detected anxiolytic effects, and fluoxetine increased anxiety-like behaviour in maze-experienced mice. Had SSRIs been required to pass the field's principal anxiety model, they would not have entered development. They were found instead by a neurochemical assay driven by a depression hypothesis.
The deepest problem
You cannot build an individual-differences science on measures that cannot rank individuals
The dot-probe is the most instructive failure in clinical psychology, and it is worth following in order.
1986
The task
The dot-probe is introduced. Europe PMC indexes no abstract — the most-cited experiment in the field is unverifiable from open sources.
2007
d = 0.45
172 studies confirm a threat bias in anxious people. The authors' own framing: "it is only an effect size of d = 0.45."
2010
d = 0.51
Attention Bias Modification "shows promise as a novel treatment."
2015
g = 0.29 ns
In clinical samples the effect is non-significant. Trim-and-fill collapses social anxiety from 0.40 to 0.14.
2025
r ≈ 0
36 variants, 9,600 participants: none demonstrated internal reliability greater than zero. A registered replication finds Bayesian null evidence — including on the bias itself.
Hold both facts together. A measure with internal consistency indistinguishable from zero produced a meta-analytic group difference of d = 0.45. These are not straightforwardly contradictory — a difference score with zero reliability can still support a group-mean contrast while being useless as an individual-difference index. But the field treated it as the latter for fifteen years: correlated it with symptoms, used it as a mediator, and built a therapy on moving it.
The same problem, four more times
Measure
Reliability / precision
What it was used for
Task fMRI
mean ICC = 0.397 90 experiments, N = 1,008
Individual-differences biomarkers — and the unreliability is specific to the regions the field wants, while control regions are excellent
Brain-wide association
Largest replicating |r| = 0.16 n = 9,500 needed for 80% power
Typical study n ≈ 25
Heartbeat counting
r = 0.16 with actual heartbeats correlates .80–.97 with the number reported
Interoceptive accuracy — the leg of the theory that a 55-study pre-registered meta-analysis then found null
Salivary cortisol
~25% of variance between countries N America d = 0.45 vs Europe d = 0.73
Stress reactivity — a between-country gap that exceeds the entire pooled anxiety effect
Cross-review
A cost signature, not a capacity signature
This is the sharpest point of contact with the stress review, and the two literatures reached it independently — neither citing the other.
Effect sizes as published · negative = impairment · tap or hover for detail
Anxiety Acute stress Clinical burnout
↔ scroll the chart sideways, or use the table view
Do not over-read this. The three literatures use non-comparable designs — anxiety findings are correlational trait self-report × task performance; acute-stress findings are within-subject experimental; burnout findings are clinical group comparisons. Method variance alone could generate part of the divergence. What the data can support: there is no evidence that anxiety produces the broad fluid-ability decrement characterising burnout, and positive meta-analytic evidence that anxiety's cost falls on speed rather than accuracy.
The convergence. The stress review reached the effort-price conclusion from bioenergetics, from the death of ego depletion (d = 0.04 and 0.06 across 59 labs), and from lateral-PFC glutamate accumulation. The anxiety literature reached the same structural conclusion from an entirely different tradition — attentional control theory, tested across 58 studies and 8,292 participants: efficiency falls, effectiveness holds, compensatory effort makes up the difference. Two fields, different methods, different decades, same distinction. That convergence is the strongest reason to take the effort-cost framework seriously as a general account rather than a local finding.
The divergence, and it matters clinically. Anxiety costs more to run; burnout runs worse. Someone who cannot face their work may be paying a raised price with intact machinery, or working with degraded machinery — and those call for opposite responses. Neither is diagnosable from self-report, because in both conditions subjective and objective impairment come apart.
One prediction that held, and is routinely misused
Attentional control theory predicts preserved output at higher processing cost. Across 58 studies it was confirmed verbatim: anxiety impaired efficiency but not effectiveness, and impaired inhibition and switching but not updating. Applied literatures in education, occupational selection and clinical outcome almost universally invoke the theory to explain effectiveness decrements — the one thing the meta-analysis found anxiety does not reliably produce. Note also that updating, the executive component arithmetic most needs, was spared: which is why working memory does not mediate the math-anxiety–performance relationship (indirect effects all non-significant, B ≤ .05).
Epidemiology
Is there an anxiety epidemic? Three claims, kept separate
Almost all public discussion of this fails by conflating three different propositions. They have three different answers.
a. Rising diagnosis and help-seeking
Enormous — and largely uninformative. UK anxiety diagnosis incidence rate ratio 3.51 (3.18–3.89) across 9.1 million young people, 2003–2018. But the decisive demonstration of what this metric measures: in April 2020 the same metric fell 47.8% below expected in a single month. Nobody thinks anxiety halved. Diagnosis rates track service access.
Cybulski 2021; Carr 2021, Lancet Public Health
b. Rising self-reported symptoms
Real but modest, and instrument-dependent. Eight of eleven studies using the General Health Questionnaire found significant increases in distress over time — and this comes from the very paper usually cited to debunk the epidemic. Baxter et al. did not find "nothing changed." They found diagnosed prevalence flat while symptom-scale distress rose. That is the prevalence inflation hypothesis, stated five years early by the people cited against it.
Baxter et al. 2014, Depression and Anxiety
c. Rising true disorder prevalence
One genuinely clean study, and it is positive. Same country, same diagnostic interview, probability samples, N = 13,561: English childhood anxiety 3.5% → 5.4%, RR 1.63 (1.37–1.93), 2004 vs 2017. But the rise was concentrated in White children (RR 1.88) with no clear change in minority-ethnic children (RR 0.85) — a pattern hard to explain by uniform true-incidence increase.
Taxiarchi et al. 2026, Br J Psychiatry
Reading. Age-standardised prevalence was flat 1990–2010 (3.8% → 4.0%), with crude case counts rising 36% — fully explained by population growth and ageing. "Epidemic" is not supportable. "A real but modest rise in youth anxiety in some high-income countries, embedded in a much larger rise in labelling" is.
The COVID surge: modelled versus observed
Approach
Estimate
What it rests on
Modelled Santomauro et al. 2021, Lancet
+25.6% (23.2–28.0)
27 anxiety studies, meta-regressed onto mobility and infection-rate proxies and extrapolated to 204 countries, most of which contributed no primary data. The narrow interval reflects uncertainty inside the model, not about whether the model is right.
Observed, same people Sun et al. 2023, BMJ
SMD 0.05 (−0.04 to 0.13) not significant
137 studies / 134 cohorts, restricted to ≥90% participant retention pre- and post-. Depression did move (0.12); women's anxiety did (0.20). General-population anxiety did not.
The modelled figure has nonetheless been baked into the official series. WHO now cites 359 million people with an anxiety disorder in 2021, against GBD 2019's published 301.4 million. A ~19% jump in two years is not new primary epidemiology — it is the COVID adjustment propagating forward. Anyone citing 359 million alongside the null same-cohort finding is citing two mutually inconsistent claims.
One epidemiological fact worth acting on. "Anxiety has the earliest onset of any disorder class" is true for the block (peak age 5.5) and false for several of its members: panic disorder's median onset is 25–27 and GAD's is 30–35. Early-intervention programmes targeting "anxiety" in childhood are targeting phobias and separation anxiety. They will not prevent GAD, whose modal onset is two decades later.
Interventions
What actually helps
Colour marks the strength of the evidence base, not the size of the effect. Comparators are named in the tooltip for every row — that is the variable that matters most.
Intervention effect sizes with 95% confidence intervals
Signs flipped so positive always favours the intervention · tap or hover for the comparator
Strong Moderate Limited Null / insufficient◌ dashed ring = no usable confidence interval published
↔ scroll the chart sideways, or use the table view
Read the comparator, not the position. Guided internet CBT sits near zero because it is an equivalence finding against face-to-face CBT — near-zero is the good result. Applied relaxation and the GAD psychotherapies look large because their comparator is treatment-as-usual, not a placebo. The only two rows measured against a credible placebo are individual CBT (0.56) and SSRIs/SNRIs (0.44) — and those are the two the evidence actually supports. Everything below MBSR has an interval crossing zero, no usable interval, or both.
What to expect
Question
Answer
How likely am I to respond to first-line treatment?
45–65% across anxiety disorders, from 403 RCTs. A third to a half of people do not respond at all — worth knowing in advance rather than concluding you are the problem.
Is online as good as in person?
Guided internet CBT is equivalent to face-to-face (g = 0.02). Unguided is substantially worse, mostly via completion: 86.7% face-to-face, 79.1% guided, 48.2% self-guided.
What happens if I stop medication?
Relapse 36.4% vs 16.4% on continuation, at up to one year. Beyond one year there is no evidence. Discontinuation symptoms occur in 31% vs 17% on placebo — a net ~15%, roughly 1 in 6.
Do CBT gains last?
Largely unstudied rather than established. At ≥12 months: GAD g = 0.22 (k = 10), social anxiety k = 3, specific phobia k = 1, OCD k = 0, and non-significant for panic disorder. Relapse was reported in 6 of 69 trials.
Negative findings
Overhyped or unsupported
Named explicitly, because the alternative is that they keep circulating.
Propranolol for performance anxiety. Near-universal among musicians and public speakers. The total evidence across all anxiety disorders is eight trials; the largest social-phobia study had 16 participants. The systematic review concludes the evidence "is insufficient to support the routine use of propranolol in the treatment of any of the anxiety disorders." It may work — nobody knows.
Attention Bias Modification. Closed question. Non-significant in clinical samples, trim-and-fill collapses it further, and a registered replication found Bayesian evidence for the null — including on the bias it was supposed to move.
Anything sold on a biomarker. Salivary cortisol panels, HRV-based "nervous system" assessments, gut-microbiome anxiety testing. Cortisol reactivity is null; HRV differences are substantially a medication effect; microbiome evidence in anxiety disorders totals three studies and 84 patients.
Apps as treatment. Against an active or placebo app the anxiety effect is g = 0.10, non-significant. Against nothing, 0.28. The difference is what expectancy and attention buy.
Ashwagandha. Four meta-analyses of overlapping trials report pooled estimates from SMD −1.55 to −6.87 — nearly seven standard deviations, larger than any psychiatric intervention ever demonstrated and physically implausible. I² of 93–98%, uniform positivity. Hepatotoxicity is verified: 25 patients, one transplant and three deaths.
CBD. The largest synthesis — 54 trials, 2,477 participants — found no significant effect on anxiety, with all-cause adverse events OR 1.75, NNTH 7. (Precision: the null is reported for cannabinoids generally; a CBD-only anxiety stratum could not be verified. There is certainly no positive signal.)
Silexan (lavender oil). Real trial evidence — but two authors are employees of the manufacturer, and the trial's active comparator paroxetine failed to separate from placebo (p = 0.10). A trial where the established drug fails but the sponsor's product succeeds should lower your confidence.
"The physiological sigh reduces anxiety." The trial compared five minutes of breathwork against five minutes of mindfulness meditation, in remote, self-selected, non-clinical volunteers. The headline positive outcomes were mood and respiratory rate, reported as p-values without effect sizes. The claim it supports is "slightly better than meditation," not "reduces anxiety."
"Some anxiety improves performance." Yerkes–Dodson is forty mice, and on meta-analysis facilitative anxiety is self-confidence under another name: cognitive anxiety r = −0.10, self-confidence r = +0.24.
"Anxiety makes people dangerous drivers." Across 24 studies, no consistent association; the one study with objective diagnosis and coroner records gave OR 0.64 and 0.69, both non-significant. What raises crash risk is benzodiazepines (OR 1.59; with alcohol 7.69).
Reassurance through testing. Across 14 RCTs and 3,828 patients, diagnostic testing to reassure produced null effects on illness worry, anxiety and symptom persistence. A negative scan does not settle the question, because the question was never about the scan.
Reference
Every key number
Filter by keyword or category. Null and non-significant findings are included deliberately — in this field they are most of the story.
Finding
Value
Source
Method
What this review corrected, and what it could not verify
Seventeen load-bearing claims were adversarially fact-checked against primary sources. No fabricated citations were found — notable, given that ten of the checked items were dated 2025–2026 and several matched down to individual confidence intervals. Five required correction.
Maddock & Smucny (2025). "Reduced total choline is the only replicated MRS finding" drops a second one — NAA was also reduced across all cortical regions. The defensible version: tCho is the only finding significant with and without outlier exclusion.
ENIGMA panic (Han et al., 2026). The quoted case-control range (d = −0.07 to −0.13) is accurate but omits the paper's largest effect: early-onset panic disorder → larger lateral ventricles at d = 0.31–0.38.
Wilson et al. (2026). The anxiety null is reported for cannabinoids generally, not CBD specifically.
Pond et al. (2025). The ABM registered replication found Bayesian null evidence on four of five measures; the depression measure was insensitive, not null-supporting.
Moncrieff et al. (2022). "Anxiet-" appears once in the readable full text, not twice. The load-bearing scope claim — that the umbrella review concerns depression only and licenses no inference about anxiety — is confirmed.
Widely quoted figures whose provenance fails
"US$1 trillion per year" for depression and anxiety. Traced to a 2016 WHO press release, and not computable from its usual citation, which reports 15-year cumulative net present values ($230bn depression + $169bn anxiety productivity gain over 2016–2030). No arithmetic path leads to $1tn/year.
WHO's "359 million with an anxiety disorder in 2021." Not new epidemiology — it propagates the modelled COVID adjustment forward from 301.4 million.
"1 in 9 → 1 in 5" UK youth mental health trend. Straddles an instrument change (diagnostic interview → screening questionnaire). Do not cite.
GAD's 6-month criterion "reduced to 3 months in DSM-5." Incorrect — it describes a rejected pre-publication draft.
Duits et al. (2015) fear-conditioning effect sizes. Closed access; no citing open-access source quotes them. Anyone citing a specific d should be asked to show the table.
von der Embse et al. (2018) test-anxiety effect sizes. The abstract contains no r, k, N, CI or heterogeneity statistic. Every specific figure attributed to it is unverified.
Stubbs et al. (2017) exercise in diagnosed anxiety disorders: the published confidence interval is arithmetically impossible — the point estimate lies outside it.
Things that do not exist — each reportable as a finding
A meta-analysis of the cortisol awakening response in anxiety disorders · a meta-analysis of basal or diurnal cortisol in GAD, panic or social anxiety · a dexamethasone-suppression meta-analysis in anxiety · any benzodiazepine-receptor PET/SPECT study in anxiety published 2009–2026 · a Cochrane review of probiotics for anxiety · a meta-analysis of oestradiol × fear extinction · an anxiety-specific equivalent of Border et al. (2019) · a meta-analysis of risk-taking in anxiety disorders · any published empirical test of the prevalence inflation hypothesis, three years after the call to test it · any well-powered study pitting avoidance against arousal as the driver of impairment · any CO₂ challenge ever performed in congenital central hypoventilation syndrome patients, 32 years after the theory that requires it was proposed.