Skip to content
Apatheia Labs
Writing

Evidence Review

Anxiety

A critical review of the evidence

Paul StephenApatheia LabsJuly 26, 2026 · 121 min read
On this page62 sections

Companion assets

AssetLink
Interactive review (HTML)/evidence-reviews/anxiety-a-critical-review-of-the-evidence/Anxiety-Critical-Review.html
Evidence table (XLSX, 134 findings)/evidence-reviews/anxiety-a-critical-review-of-the-evidence/Anxiety-Evidence-Table.xlsx
Practical brief/evidence-reviews/anxiety-a-critical-review-of-the-evidence/Anxiety-Practical-Brief.md

Companion review: Stress, Energy and the Capacity to Function.


Executive summary

The stress review found a field where the folk model was wrong but a better mechanistic story was waiting underneath. Anxiety is different, and harder. Here the folk model is also wrong — but what replaces it is largely a set of well-powered nulls, demonstrated confounds and outright reversals. The honest headline is that anxiety research has spent forty years failing to find a pathophysiology, and the most important recent developments are negative.

That is not a nihilistic conclusion. Two treatments work, their effect sizes are known once you insist on honest comparators, and several widely promoted alternatives can be confidently set aside. But anyone expecting a clean causal account of what anxiety is will not find one in the evidence.

Ten things the evidence actually supports:

  1. The neuroanatomical basis of the fear/anxiety distinction has evaporated in humans. The textbook model — amygdala for phasic fear, BNST for sustained anxiety — was directly tested in a harmonised mega-analysis of 295 adults. The two regions responded to threat in a statistically indistinguishable way, with positive Bayesian evidence of equivalence, and the responses ran opposite to the model's predictions (Didier et al., 2026).

  2. "Amygdala hyperactivity" is not established as a marker of anxiety disorder. It rests on a 2007 coordinate-based meta-analysis that reports no effect sizes, built from a measurement modality whose test–retest reliability averages ICC = 0.40 (Elliott et al., 2020) and where the largest replicating brain-wide association is |r| = 0.16 (Marek et al., 2022). At consortium scale, ENIGMA found no effect of GAD on brain structure at all, and structural effects in panic and social anxiety are d ≈ 0.07–0.13.

  3. The interoceptive-accuracy account is empirically dead, and the instrument that generated it is broken. A pre-registered meta-analysis of 55 studies found no association between cardiac interoceptive accuracy and anxiety (Adams et al., 2022). The heartbeat-counting task correlates with actual heartbeats at r = 0.16 and with the raw number a person reports at .80–.97 — a cardiac recording is barely needed.

  4. Reduced heart-rate variability in anxiety is substantially a medication effect. In NESDA, unmedicated anxious patients did not differ from controls; starting antidepressants lowers RSA and stopping reverses it. And the field's most-quoted HRV effect size, g = −0.45, is a post-exclusion figure: the primary estimate was g = −0.70 (−1.45, −0.05), p = 0.07, I² = 98.2%, and the deleted "outlier" was the one study large enough to test the medication confound.

  5. HPA-axis hyperactivity in anxiety is null, and where it has a sign it points down. Cortisol stress reactivity: AUCi = −0.101, p = 0.28. Cortisol awakening response: r = .03, p = .13. Diurnal slope: anxiety was the only one of twelve outcomes with a negative coefficient. For calibration, ~25% of variance in laboratory cortisol responses is attributable to systematic differences between countries — a between-country gap that exceeds the entire pooled anxiety effect.

  6. The dot-probe — the measure on which two decades of attentional-bias theory and an entire intervention were built — has an internal reliability of zero. Across 36 variants and 9,600 participants, "none of the 36 versions demonstrated internal reliability greater than zero" (Xu et al., 2025). Attention Bias Modification correspondingly failed: from d = 0.51 in 2010 to a non-significant g = 0.29 in clinical samples, to a registered replication finding Bayesian evidence for the null on four of five measures — including on the bias itself.

  7. Anxiety's cognitive signature is a cost signature, not a capacity signature — and this is the sharpest point of contact with the stress review. Across 58 studies and 8,292 participants, anxiety impaired attentional control efficiency but not effectiveness, and impaired inhibition and switching but not updating (Shi et al., 2019). Anxious people pay more to reach the same output. Clinical burnout, by contrast, produces a genuine ability-level decrement. Two literatures, developed independently, converge on the same distinction the stress review reached from bioenergetics: the price of effort rises before the capacity for it falls.

  8. Yerkes–Dodson does not survive contact with its source. The 1908 paper used roughly forty Japanese dancing mice in a brightness discrimination punished by electric shock, with two mice per data point in the only condition that produced an inverted-U; the easy condition was monotonic; the paper contains no inferential statistics and never uses the word "arousal." There is no good evidence for an inverted-U between self-reported anxiety and cognitive performance.

  9. There is no anxiety epidemic, but there is one clean positive finding inside a much larger labelling effect. Age-standardised prevalence was flat 1990–2010 (3.8% → 4.0%). UK anxiety diagnosis rates rose with an incidence rate ratio of 3.51 — the same metric that fell 47.8% in a single month in April 2020, which tells you what it measures. Against that, one methodologically clean same-instrument comparison found English childhood anxiety rising 3.5% → 5.4%, RR 1.63. And the modelled 25.6% COVID surge is contradicted by 134 same-participant cohorts showing SMD 0.05 (−0.04 to 0.13), non-significant.

  10. Two treatments work, and both are roughly half as large as advertised. In social anxiety, individual CBT measures −1.19 against a waitlist and −0.56 against a psychological placebo; SSRIs measure −0.91 against a waitlist and −0.44 against a pill placebo. They were the only two intervention classes that beat their appropriate comparator. Response to first-line treatment is 45–65%, meaning a third to a half of patients do not respond at all.

And one structural conclusion that runs underneath all four domains. The genetic correlation between anxiety and depression is rg = 0.91 in the largest GWAS to date. Quantitative structural models place GAD closer to major depression than to panic or the phobias. Intolerance of uncertainty is transdiagnostic. The general psychopathology factor performs no better than a raw count of comorbid diagnoses. The category "anxiety disorders" as the manuals draw it is not obviously carving nature at a joint — and a good deal of the failure to find a pathophysiology may follow from looking for one under a heading that does not name a natural kind.


How to read this review

The single most useful heuristic in this literature

An effect size without a named comparator is not a finding. In social anxiety disorder, individual CBT is −1.19 against a waitlist and −0.56 against a psychological placebo. SSRIs are −0.91 against a waitlist and −0.44 against a pill placebo. Both lose approximately half their apparent magnitude when the control is made honest.

More than 80% of CBT anxiety trials use waitlist controls, and only 17.4% are high quality (Cuijpers et al., 2016). Any effect size quoted in this field without its comparator should be assumed to be the waitlist number, and mentally halved.

Three structural problems

Waitlist inflation. In a network meta-analysis of 49 RCTs, the odds of response for no treatment over waiting list was 2.9 (1.3–5.7) — people given nothing did better than people told to wait. (Adversarial note: this most-cited finding is fragile. The authors state evidence quality "was less than ideal," and the significant differences did not survive post-hoc adjustment for publication bias. It is a strong hypothesis, not a settled parameter; the convergent evidence above is what carries the claim.)

Publication bias. Across 74 FDA-registered antidepressant trials, the published literature suggested 94% were positive; the FDA data showed 51%, with effect-size inflation of 32% overall (Turner et al., 2008). Preregistration has improved this for drugs. It has not improved for digital therapeutics: across 176 mental-health-app RCTs from 2011–2023, methodological quality showed minimal improvement over twelve years.

Measurement. This is the problem specific to anxiety, and it is worse here than in the stress literature. The dot-probe has zero internal reliability. Task-fMRI has a mean ICC of 0.40. Heartbeat counting correlates r = 0.16 with actual heartbeats. Several of the field's central constructs were operationalised by measures that cannot in principle support the individual-differences claims made from them.

Conventions

Claims are marked [contested] where the literature genuinely disagrees and [unverified] where a figure is widely repeated but could not be traced to a primary source. Appendix C lists every unverified claim, and it is long on purpose. Confidence intervals are omitted rather than reconstructed where they could not be verified.

What this review corrected

Seventeen load-bearing claims were adversarially fact-checked against primary sources. No fabricated citations were found — notable, given that ten of the checked items were dated 2025–2026. Five required correction:

  • Maddock & Smucny (2025): the claim that reduced total choline is the only replicated MRS finding drops a second one — NAA was also reduced across all cortical regions. The defensible version is that tCho is the only finding significant with and without outlier exclusion.
  • ENIGMA panic (Han et al., 2026): the quoted case-control range (d = −0.07 to −0.13) is accurate but omits the paper's largest effect — early-onset panic disorder was associated with larger lateral ventricles at d = 0.31–0.38. Section 1.3 now states both.
  • Wilson et al. (2026): the anxiety null is reported for cannabinoids generally, not CBD specifically. Narrowed in §4.9.
  • Pond et al. (2025): the ABM registered replication found Bayesian null evidence on four of five measures; the depression measure was insensitive, not null-supporting. Corrected in §2.2.
  • Moncrieff et al. (2022): the string "anxiet-" appears once in the readable full text, not twice. The load-bearing scope claim — that the umbrella review concerns depression only — is confirmed.

Three items in the commissioning briefs were themselves miscited and are corrected here: Baxter et al.'s "Challenging the myth of an 'epidemic'" is in Depression and Anxiety, not Psychological Medicine; Schmidt, Zvolensky & Maner (2006) is in Journal of Psychiatric Research; Hancock & Ganey (2003) is in Journal of Human Performance in Extreme Environments.

Part I — Mechanisms: forty years of looking for a pathophysiology

1.1 Fear versus anxiety: the dissociation that failed its best test

The textbook model is clean. Phasic fear — a response to imminent, identifiable threat — is mediated by the central nucleus of the amygdala. Sustained anxiety — apprehension about diffuse or uncertain threat — is mediated by the bed nucleus of the stria terminalis (BNST) via corticotropin-releasing factor. The distinction comes from Davis, Walker, Miles & Grillon (2010) and was imported wholesale into the NIMH Research Domain Criteria framework.

Shackman & Fox (2016) argued it was already untenable: both regions "are sensitive to a range of aversive challenges… both show phasic responses to short-lived threat; and both show heightened activity during sustained exposure to diffusely threatening contexts."

The decisive human test has now been run, and it went badly for the model. Didier, Grogans and colleagues (2026) report a harmonised mega-analysis of fMRI from 295 adults on a threat-anticipation paradigm. Verbatim: "Contrary to popular double-dissociation models, results demonstrated that the Ce responds to temporally uncertain threat and the BST responds to certain threat. In direct comparisons, the two regions showed statistically indistinguishable responses, with strong Bayesian evidence of regional equivalence." Frontocortical regions, not the extended amygdala, carried the preferential uncertain-threat response.

Note what this is: not a failure to reject the null, but a positive demonstration of equivalence via two-one-sided-tests. [contested: a reply by Wang, Tseng and colleagues runs in the same Nature Reviews Neuroscience exchange — the debate is live, not closed.]

What survives. The rodent dissociation may be real at the circuit level. The fear/anxiety distinction may remain valid psychologically and clinically — the phenomenology of a panic attack and of chronic worry genuinely differ. What has evaporated is its human neuroanatomical warrant. A distinction that was being used to justify a mechanistic taxonomy is back to being a description.

1.2 LeDoux's reframe, and the evidence that actually supports it

LeDoux & Pine (2016) distinguish two systems: (1) defensive survival circuits producing behavioural and physiological responses, and (2) conscious feeling states reported verbally. Conflating them, they argue, "has impeded progress." On this account "fear circuit" is a category error — amygdala circuits are threat-detection and defence circuits, not fear circuits, and the feeling is a separate cortical accomplishment.

Fanselow & Pennington (2018) reply that the framework "carries the frightening implication that we ought to reduce the study of fear to subjective report," and defend a central fear generator with evolutionarily conserved defensive function.

The strongest empirical argument for the two-system view is rarely made, and it deserves to be. Three independent findings show clean dissociation between circuit engagement and subjective outcome:

  • A CRF1 antagonist significantly enhanced fear-potentiated-startle inhibition in PTSD patients "independent of clinical effects" — while the parent trial failed (Jovanovic et al., 2020).
  • Pexacerfont produced no change in anxiety "despite drug levels in CSF that predict close to 90% central CRH1 receptor occupancy" (Kwako et al., 2015).
  • Hypoventilation therapy in panic disorder normalised PCO₂ and respiratory rate, but "improvements… for symptom severity and anxiety did not differ between TX and WL" (Tunnell et al., 2021).

In each case the physiology moved and the experience did not. Whatever one thinks of the philosophy, this is a practical warning: normalising an anxiety-associated physiological measure does not reliably normalise anxiety. That has direct implications for the wearables and biofeedback market, which sells the inference in the opposite direction.

1.3 The hyperactive amygdala, stress-tested

The claim. Etkin & Wager (2007) reported that patients with PTSD, social anxiety disorder or specific phobia "consistently showed greater activity than matched comparison subjects in the amygdala and insula." This is the single most-cited functional finding in anxiety neuroscience.

Three problems, in ascending order of severity.

First, the source reports no effect sizes. It is a coordinate-based meta-analysis. It tells you where activations converge, not how large or how reliable they are.

Second, the newer and larger meta-analysis does not find it. Chavanne & Robinson (2021) screened 3,433 records and included 181 articles (induced anxiety N = 693; pathological anxiety N = 2,554 patients / 2,348 controls). The convergence between induced and pathological anxiety was bilateral insula and cingulate/medial prefrontal cortex. The amygdala does not appear in the reported overlap. [unverified: whether the amygdala appeared in the pathological-anxiety-only map — full text not obtained.]

Third, and decisively, the measurement cannot support the inference.

FindingValueSource
Task-fMRI test–retest reliabilitymean ICC = 0.397 (90 experiments, N = 1,008); 11 tasks ICC 0.067–0.485Elliott et al., 2020
Amygdala/sgACC response to faces, 14-day retestRobust group activation, low within-subject reliability — while the control FFA region was excellentNord et al., 2017
Median univariate brain-wide association|r| = 0.01; largest out-of-sample-replicating |r| = 0.16Marek et al., 2022
Sample needed for 80% power on the top 1% of effectsn = 9,500 (Bonferroni)Marek et al., 2022

Nord et al.'s result is the sharpest: the unreliability is specific to the regions the field wants as biomarkers. It is not a general noise problem.

What survives at consortium scale. The ENIGMA-Anxiety working group has now assembled the largest structural datasets in the field:

  • GAD (1,020 cases / 2,999 controls): verbatim, "The main analysis showed no effect of GAD on brain structure, nor interactions involving GAD, age, or sex."
  • Social anxiety (1,115 / 2,775): bilateral putamen smaller, left d = −0.077, right d = −0.104.
  • Panic (N = 4,924): case-control effects d = −0.07 to −0.13. But note the paper's largest finding, which is not a case-control contrast: early-onset panic disorder (≤21 years) was associated with larger lateral ventricles at d = 0.31–0.38 — three to five times anything in the case-control range.

Verdict. Group-average amygdala responses to threat are robust and real. Amygdala hyperactivity as an individual-differences marker of anxiety disorder is not established, the best-powered structural data do not support it, and the modality's reliability sits below what individual-differences research requires. The early-onset ventricular finding is the one structural signal large enough to be worth pursuing, and it points at neurodevelopment rather than at a threat circuit.

1.4 Interoception: the accuracy leg is dead, the instrument is broken

Paulus & Stein (2006) proposed that anxiety-prone individuals show "augmented detection of the difference between the observed and expected body state," with anterior insula central. Barrett & Simmons (2015) formalised it as interoceptive inference. Garfinkel et al. (2015) established that interoceptive accuracy (objective performance), sensibility (self-reported attentiveness) and awareness (metacognitive correspondence) are dissociable.

The critical test has been run and it is null. Adams et al. (2022), pre-registered, 55 studies: "Overall, we found no evidence for an association between cardiac interoceptive accuracy and anxiety, with none of the factors examined moderating this finding."

And the instrument that produced the original literature does not measure what it claims. In N = 572, heartbeat-counting accuracy scores above 95% reflect under-reports; the correlation between actual and reported heartbeats is r = 0.16; and scores are structurally bound to resting heart rate (Zamariola et al., 2018). A follow-up found heartbeat-counting scores correlate .80–.97 with the raw count of reported heartbeats — meaning the cardiac recording contributes almost nothing. [contested: Zimprich et al. (2020) argue these are ratio-variable artefacts and "not as serious as they might appear."]

What remains is the self-report leg, and it may be partly tautological. Across 71 studies and 20 subscales, anxiety is associated with increased negative evaluation of, sensitivity to and negative attention toward bodily signals — but not with using bodily signals to inform emotion (Clemente et al., 2024). The authors themselves flag "the overlap between anxiety and interoception questionnaires." A questionnaire that asks whether bodily sensations worry you will correlate with anxiety by construction.

Neurally, a pre-registered meta-analysis of 11 studies found reduced insula activation in major depression but "no significant effects were observed for anxiety disorders or the combined sample" (Bednarek et al., 2026).

1.5 Autonomic markers, and a confound that accounts for the effect

The headline claim. Chalmers et al. (2014) meta-analysed 36 articles (2,086 patients / 2,294 controls) and is universally cited for reduced HRV in anxiety disorders. High-frequency HRV across all disorders: g = −0.29 (−0.41, −0.17), I² = 35.8%.

But the time-domain figure everyone quotes is a post-exclusion number, and this matters. The primary estimate across 20 datasets was g = −0.70 (−1.45, −0.05), p = 0.07, I² = 98.2% — non-significant with essentially total heterogeneity. The frequently quoted g = −0.45 (−0.57, −0.33), I² = 2.5% appears only after removing a "clear outlier."

That outlier was Licht et al. (2009) — the NESDA cohort, n = 2,059, comprising more than half of all pooled patients, and the single study designed to test the medication confound. Chalmers' own table records it as "Effect disappears when controlling for psychotropic use."

The medication confound is causal, not correlational:

  • Licht et al. (2009): "Nonmedicated anxious subjects did not differ from controls."
  • Licht et al. (2010), N = 2,114: starting antidepressants lowers respiratory sinus arrhythmia; stopping reverses it.
  • ELSA-Brasil (Kemp et al., 2014), N = 15,105, propensity-weighted: only GAD survived, "though small," against tricyclic effects of d = 0.72–0.81.

A 2025 umbrella review of 442 primary studies grades no anxiety disorder above Class IV evidence; GAD's HF, RMSSD and LF/HF are all non-significant.

And the construct itself is contested by the physiologists. An international Expert Recommendation in Nature Reviews Cardiology (Menuet et al., 2025) states that respiratory sinus arrhythmia amplitude "should not be misconstrued as a measure of vagal tone." Decisively for case-control work: HF-HRV correlates with respiration rate in patients but not in controls (Quintana et al., 2016) — the respiration confound is differential, so it can manufacture the group difference outright.

Klein's false suffocation alarm deserves a paragraph because it is the most elegant theory in this section and it fails. Its two cleanest predictions are wrong: the direct chemoreflex slope is normal in panic disorder, and CO₂ hypersensitivity failed as a familial risk marker in Klein's own collaborators' 142-offspring study. Specificity is a gradient, not a boundary (panic vs healthy OR 11.5; panic vs other psychiatric disorders OR 3.52). The sharpest point is internal: Klein's own arterial data showed "High Pco₂ levels and low pH levels virtually precluded developing panic" — which he conceded "does not fit a simple carbon dioxide hypersensitivity model" — and rescued with an unmeasured, fluctuating alarm threshold, then read the same low-pCO₂ state as a risk marker in patients and a protective factor in pregnancy. As deployed, the theory is unfalsifiable.

1.6 The HPA axis: null, and where it has a sign it points down

This is the sharpest contrast with the stress literature, and it is under-appreciated.

MeasureEstimateSource
Cortisol stress reactivity (AUCi), anxiety−0.101 (−0.285, 0.082), p = 0.28Zorn et al., 2017 (9 studies, n = 732)
Cortisol stress output (AUCg)−0.031, p = 0.87Zorn et al., 2017
Cortisol awakening responser = .03, p = .13 (k = 59)Boggero et al., 2017
Diurnal sloper = −.084 (−.173, .006), p = .066Adam et al., 2017
Depression, for contrastd = 0.60 → d = 0.33 in methodologically adequate studiesStetler & Miller, 2011

The single positive result in Zorn et al. is one uncorrected sex subgroup among more than forty tests, from a pool in which two of nine studies are in 9–10-year-olds and one failed to move cortisol in controls at all. The published abstract never states the overall null.

A calibration figure worth keeping. Across 237 Trier Social Stress Test studies and 8,487 participants, "~25% of this variability is actually attributable to systematic differences between countries" — North America d = 0.45 versus Europe d = 0.73 (Miller & Kirschbaum, 2019). A between-country difference of Δd ≈ 0.28 exceeds the entire pooled anxiety effect.

And large parts of this literature simply do not exist. There is no meta-analysis of the cortisol awakening response in anxiety disorders, none of basal or diurnal cortisol in GAD, panic or social anxiety, none of the dexamethasone suppression test in anxiety, and none of the dex/CRH test for any condition. The entire dex/CRH literature in anxiety amounts to roughly 66–84 patients ever, with zero GAD and zero social anxiety studies.

The practical upshot: cortisol is not a useful biomarker in anxiety, and salivary cortisol panels sold as anxiety assessments have no evidential basis. This mirrors the stress review's finding on "adrenal fatigue," from a different direction.

1.7 Neurochemistry, handled with scope discipline

GABA. The intuition is strong — benzodiazepines work, they act at the GABA-A receptor, therefore anxiety involves GABA deficiency. The evidence does not deliver it. The human benzodiazepine-receptor imaging literature in anxiety is five studies, n = 7–15 per group, published 1997–2008, disagreeing on direction and region. Verified negative: zero such studies published 2009–2026, and no meta-analysis exists.

The first and only ¹H-MRS meta-analysis in anxiety disorders (Maddock & Smucny, 2025; 25 datasets, 370 patients) reports GABA g = −0.26, p = 0.52, I² = 66%, from 3 datasets and 42 patients. Glutamate and Glx are null in every region and pooling. Two metabolite findings did reach significance — reduced total choline and reduced NAA across cortical regions — but tCho is the only one significant with and without outlier exclusion, and choline is a membrane marker unrelated to the GABA hypothesis.

The inference from drug efficacy to pathophysiology is the ex juvantibus fallacy, memorably demolished by Lacasse & Leo (2005): "the fact that aspirin cures headaches does not prove that headaches are due to low levels of aspirin in the brain."

Serotonin — and a scope error worth naming. Moncrieff et al.'s (2022) umbrella review of the serotonin hypothesis is frequently invoked in discussions of anxiety. It is a review of 17 studies about depression. Full-text verification confirms the string "anxiet-" appears once in the readable text, inside a journal name in the reference list. It licenses no inference about anxiety. Its methodological critique — directional ambiguity of binding data, medication confounding, underpowering — transfers cleanly, to a literature with samples an order of magnitude smaller.

(The rebuttals are substantive — Jauhar and 35 co-signatories showed Moncrieff mischaracterised 5-HT1A receptors as exclusively presynaptic autoreceptors, reversing a key directional inference. But note the boomerang: having criticised Moncrieff for discussing efficacy in a review presenting no efficacy data, their closing sentence is "The proven efficacy of SSRIs… lends credibility to this position" — the same fallacy, in the other direction.)

Directly in anxiety, the serotonin evidence points opposite to a deficit. Frick et al. (2015) conclude social anxiety disorder is characterised by an "overactive presynaptic serotonin system, with increased serotonin synthesis and transporter availability." And acute tryptophan depletion does not reliably provoke anxiety: across 21 studies, "unlike in depression, ATD does not indicate vulnerability to develop an anxiety disorder" — in GAD, depletion did nothing despite a 92% reduction in the free tryptophan ratio.

Efficacy is real and aetiologically uninformative. Drug-minus-placebo across anxiety disorders is g = 0.26–0.39, against a placebo pre-post improvement of g = 0.70–1.10 (Sugarman et al., 2017). In panic disorder, the Cochrane network meta-analysis ranks SSRIs fifth of five classes, with no class differing from another. Drugs that work do not tell you what is broken.

1.8 Genetics: what actually scales

Twin heritability is lower than commonly quoted: panic 0.48 (0.41–0.54), GAD 0.316 (0.24–0.39) — the latter resting on two twin studies (Hettema et al., 2001). The familiar "~50%" is an upper bound from a measurement-error correction applied to one sample.

The current ground truth is Strom et al. (2026), the PGC Anxiety Working Group: 122,341 cases / 729,881 controls, 58 genome-wide significant SNPs (51 replicated), SNP-heritability 10.1% on the liability scale — which the authors note "captures approximately one-quarter of the broad-sense heritability from twin studies." The strongest cell-type association was GABAergic neuroblasts (P = 3.24 × 10⁻⁸), which is the one point at which the GABA story finds independent support, from a completely different method.

Polygenic scores explain 2.27% of variance in European-ancestry samples and 0.54% (p = 0.051, non-significant) in African-ancestry samples — a portability failure that should temper any clinical-prediction ambition.

Note the correction to an influential earlier paper. Purves et al. (2020) reported SNP-heritability of 26–31% and titled their paper around "a major role for common genetic variation." At larger scale that estimate does not hold; the modern range is 5–10%. And Otowa et al.'s (2016) two reported loci appear in no modern hit list.

The candidate-gene cautionary tale, with a scope caveat. 5-HTTLPR × stressful life events gives OR = 1.01 (0.94–1.10) across N = 14,250; an individual-level analysis of N = 38,802 found "no subgroups or variable definitions for which an interaction… was statistically significant." Border et al. (2019), N = 621,214, found exactly one of 640 tests passing a liberal threshold — in the protective direction. But Border tested depression outcomes; anxiety was not among them. There is no anxiety equivalent of Border et al. That is a real gap, and reviews routinely paper over it. The nearest anxiety-specific meta-analysis was positive for TMEM132D and COMT — variants that then failed to appear in any modern GWAS.

1.9 Inflammation and the gut: a depression-comorbidity artefact

The canonical GAD estimate rests on a single marker: CRP, d = 0.38 (0.06–0.69), I² = 75% (Costello et al., 2019). A frequently cited overall effect for "inflammation in anxiety" (g = −0.39) is confined by its own moderation analysis to PTSD (g = −0.68).

The strongest observational finding is a reversal. In up to 144,890 UK Biobank participants, "the association of CRP with anxiety symptoms switched its valence after adjusting for depressive symptoms (OR = 0.98; 95% CI 0.97–0.99)." And Mendelian randomisation gives CRP → anxiety OR = 0.87 (0.80–0.95) — protective (Ye et al., 2021), with an independent MR study finding null. Neither supports a causal risk-increasing effect.

Probiotics. There is no Cochrane review (verified: zero hits). The cleanest statement of the translation gap comes from a single paper using identical methods on both literatures: animals g = −0.47 (p = 0.004); humans g = −0.12 (p = 0.151), null (Reis et al., 2018). The largest recent synthesis (72 RCTs) reports an overall SMD of −0.44, but its only genuinely clinical subgroup is null (k = 4, p = 0.06), Lactobacillus alone is null (p = 0.54) despite carrying the entire preclinical story, and the effect declines monotonically with trial duration.

Microbiome composition in anxiety disorders is three studies and 84 patients — from a review that also reports 57 of 59 studies (96.6%) across all disorders found significant differences. A 96.6% positivity rate is a forking-paths signature, not a discovery rate.

1.10 Sex differences: the ratio is 1.6–1.8, not 2:1

Three independent methods converge. WHO World Mental Health surveys (N ≈ 73,000): any anxiety disorder OR 1.7 (1.6–1.8) lifetime. US CPES (N = 20,013): lifetime 33.3% women vs 22.0% men, a crude prevalence ratio of 1.51 — the frequently quoted 1.70/1.79 are adjusted odds ratios, which overstate the prevalence ratio for an outcome this common. GBD 2019: 1.64:1.

The ratio is not uniform, and that is informative. It is largest for PTSD (OR 2.6) — which is no longer an anxiety disorder in DSM-5, so its inclusion in older estimates inflates the headline. It is near parity for social anxiety (1.3, and non-significant in the US data). Early-onset OCD skews male.

A hard constraint the field under-uses: in the WHO surveys the gender × birth-cohort interaction is null for every anxiety disorder, while the depression gap narrowed and the alcohol gap widened in step with changing gender roles. Whatever drives the anxiety gap has not moved.

The hormonal explanation runs the wrong way in the best-powered test. A randomised placebo-controlled 2×2 with N = 116 found "the estradiol groups showed heightened SCR responses to the previously extinguished stimulus, i.e., impaired extinction recall" (Kaczmarczyk et al., 2024) — the opposite of the founding claim, which rests on median-split designs at n = 17–18 per cell. No meta-analysis of oestradiol × fear extinction exists.

Measurement artefact is close to ruled out. Full measurement invariance of the GAD-7 across sex in N = 165,872 (largest ΔCFI = −0.006); scalar invariance on HADS-A with the female elevation remaining; and two direct experimental tests of "men under-report" that both failed. Clinician diagnostic bias for anxiety is speculation — a 475-therapist randomised vignette study found "no significant associations between gender and diagnostic decisions." [unverified: no direct differential-item-functioning study on core diagnostic instruments (CIDI modules) by sex was located despite dedicated searching — a genuine gap.]

1.11 Animal models and the drug-development failure

Since the benzodiazepines, essentially nothing new has worked. Every mechanism that succeeded in rodents has failed in humans — usually with assay sensitivity intact, which rules out the easy excuse that the trials were underpowered:

  • Pexacerfont in GAD: response 42% vs placebo 42%, while escitalopram separated at weeks 1, 2, 3, 6 and 8 in the same trial.
  • Verucerfont in PTSD failed; its authors concluded CRF1 antagonists "lack efficacy as monotherapy agents for these conditions."
  • Aprepitant failed in five depression trials with paroxetine significant in all three that included it, and PET-confirmed sustained receptor blockade.
  • Tebideutorexant (orexin antagonist) failed in 2025: N = 222, HAM-A difference 0.23, p = 0.59.

The elevated plus maze — the workhorse assay — has been formally audited and it fails. Across 814 studies and 25 compounds, pre-registered: "only two out of 17 commonly used test measures reliably detected effects of anxiolytic compounds"; there is "considerable between-study variation in size and even direction"; the authors conclude the findings "cast serious doubt on both construct and predictive validity" (Rosso et al., 2022).

Worse, its benzodiazepine sensitivity is abolished by a single prior five-minute exposure, restored by moving the animal to a different room, and manufacturable by a post-trial memory enhancer. Whether a compound is classified as anxiolytic depends on how stressed the animals were before entering the maze.

The most telling fact in this section: fluoxetine significantly increased anxiety-like behaviour in maze-experienced mice. Had SSRIs been required to pass the elevated plus maze, they would not have entered development. They were found instead by a neurochemical uptake assay driven by a depression hypothesis. The field's principal anxiolytics were discovered despite its principal anxiety model, not because of it.

Across 152 prospectively replicated mouse-behaviour results in three laboratories, the probability that a single-lab statistically significant discovery is non-replicable was 59.6%.

1.12 The pattern

Read Part I as a whole and the same arc appears in every subsection: a striking small-sample finding, a mechanistic story built on it, and then one of three outcomes —

  • a well-powered null (interoceptive accuracy, HPA reactivity, ENIGMA structural findings, probiotics in clinical samples, GABA on MRS);
  • a demonstrated confound that fully accounts for the effect (antidepressants for HRV, depression comorbidity for CRP, respiration for HF-HRV, medication for the "outlier" that was excluded);
  • a direct reversal (serotonin up in social anxiety, oestradiol impairing extinction, CRP protective, Ce and BST responding to the opposite threat types).

The few things that scale are non-specific: ~10% SNP heritability, rg ≈ 0.9 with depression, GABAergic cell-type enrichment, and modest real drug efficacy. They describe a dimension shared with depression rather than a discrete pathophysiology of anxiety.

This is worth stating plainly because the alternative reading — that we simply need better instruments — is not obviously right. It is at least as likely that the category is wrong. Section 3.7 makes that case from a different direction.

Part II — Cognition, decision-making and function

2.1 The dot-probe: a measurement reckoning

This is the most instructive failure in clinical psychology, and it is worth telling in order.

The claim. Anxious people preferentially orient attention toward threat. MacLeod, Mathews & Tata (1986) introduced the dot-probe to measure it. (A telling detail: Europe PMC indexes no abstract for the single most-cited experiment in the field; its N and effect size are unverifiable from open sources.)

The meta-analysis. Bar-Haim et al. (2007): 172 studies, N = 2,263 anxious and 1,768 non-anxious, d = 0.45. The authors' own framing was already deflationary — "it is only an effect size of d = 0.45."

The reliability problem, in three escalating steps.

  • Schmukle (2005), verbatim: internal consistency and one-week retest estimates "lead to the conclusion that the dot probe task is a completely unreliable measure of attentional allocation in non-clinical samples."
  • Rodebaugh et al. (2016): "Our own data replicate previous findings of poor reliability for traditionally used scores." They note the benchmark — mediation analyses require reliabilities of .90 or greater.
  • The decisive datum. Xu et al. (2025) tested thirty-six variants of the emotional dot-probe (faces / scenes / snakes-spiders × SOA 100/500/900 ms × orientation × trial type) in N = 9,600. Verbatim: "none of the 36 versions demonstrated internal reliability greater than zero. Reliability was similarly poor in anxious participants."

Hold both facts together. A measure with internal consistency indistinguishable from zero produced a meta-analytic group difference of d = 0.45. These are not straightforwardly contradictory — a difference score with zero reliability can still support a group-mean contrast while being useless as an individual-difference index. But the field treated it as the latter for fifteen years, correlated it with symptoms, used it as a mediator, and built a therapy on moving it.

The general lesson, which generalises well beyond anxiety: an experimental effect that replicates at the group level tells you nothing about whether the measure can rank individuals. Reliability and validity are separate questions, and the second is worthless without the first.

2.2 Bias modification: what failed and what survived

Attention Bias Modification. The trajectory is brutal and instructive.

StageFinding
2010 meta-analysis12 studies, 467 participants, d = 0.51 — "shows promise as a novel treatment"
2015 meta-analysis, all samples41 comparisons, g = 0.37 (0.20–0.54), I² = 72.7%
2015, clinical samplesg = 0.28 (0.01–0.55)
2015, general anxiety, clinicalg = 0.29 (−0.14 to 0.74) — not significant
2015, social anxiety, clinicalg = 0.32 (−0.09 to 0.74) — not significant
2015, social anxiety after trim-and-fill0.40 → 0.14 (−0.22 to 0.51), not significant
2025 registered replicationN = 104 high worriers, Bayesian evidence for the null on four of five measures — including on the bias itself

(Correction applied: the registered replication found the depression measure insensitive*, not null-supporting. Null evidence held for the composite index, PSWQ, STAI-T and the bias measure.)*

The theoretical response has been a rescue rather than a replication: Mogg & Bradley concede the "disappointing effectiveness of ABM-threat-avoidance training" and reconceive ABM as goal-directed cognitive skill training rather than bias correction.

Interpretation bias modification is the one member of the family with a defensible signal. The best-designed synthesis (Fodor et al., 2020; 85 trials, 65 in anxiety, n = 3,897) found IBM beat waitlist at SMD −0.55 (−0.91 to −0.19) and sham at SMD −0.30 (−0.50 to −0.10). ABM showed benefit only in post-hoc sensitivity analyses excluding PTSD trials.

Two caveats the authors state outright: "Prediction intervals for all findings were large, including an SMD of 0," and "only four randomised controlled trials had low risk of bias on all six domains."

Verdict: IBM is worth continued investigation at roughly −0.30 against sham, with a prediction interval spanning zero. Transfer to clinical symptoms is not established. ABM should be regarded as a closed question.

2.3 Worry: three theories, none winning

Borkovec's cognitive-avoidance model — worry as verbal-linguistic avoidance of aversive somatic imagery — is in the worst empirical shape of the three, and the problems start in the founding studies.

  • The canonical demonstration reports in its own abstract that the worry group "showed significantly greater subjective fear" than the neutral condition. A theory of emotional avoidance resting on a manipulation that increased fear.
  • The cardiovascular suppression effect occurred "only on the first image presentation," and in only one of three worry conditions.
  • The suppression effect is a baseline artefact: "When a pre-worry resting baseline is used as the comparison point, there is no muting effect of worry on reactivity to fear stimuli" (Newman & Llera, 2011).
  • Direct disconfirmations in GAD samples: skin conductance increased with worry, with "no suggestion that worry suppressed affective responding."
  • Meta-analytically decisive (Ottaviani et al., 2016; k = 60): perseverative cognition raises systolic BP g = .45, diastolic g = .51, heart rate g = .28, cortisol g = .36, and lowers HRV. Every sign is opposite to autonomic suppression.

Contrast Avoidance (Newman & Llera, 2011) — worry as a strategy for maintaining negative affect to avoid a sharp increase in it — reinterprets the same data coherently and has the best ecological-momentary support. But its causal-experimental base is five unpreregistered lab studies with Ns of 68–96 reporting d = 0.84–1.54. At roughly twelve participants per cell, those should be treated as upper bounds inflated by sampling error. The one independent psychophysiological test reports the data "partly support" the model, in healthy participants.

Intolerance of uncertainty has by far the most data and the weakest specificity.

  • IU × GAD r = .57; IU × depression r = .53; IU × OCD r = .50. The GAD–depression difference was not significant.
  • Across 181 studies, N = 52,402, 335 effect sizes: r = 0.51 (0.50–0.52) — with the authors' own conclusion being "definitive evidence for the transdiagnostic nature of IU."
  • The entire causal foundation is a single study with n = 21 per cell, with no independent direct replication located in 26 years.
  • Most damaging: in latent-variable modelling, all four "vulnerability" constructs loaded on the anxiety facet of neuroticism, and "after accounting for the contribution of neuroticism facets, intolerance of uncertainty and experiential avoidance were not uniquely associated with any disorders" (Naragon-Gainey & Watson, 2018).
  • And targeting it clinically: vs passive control g = −0.94; vs active controls, pooled effects were not significant.

Which wins? None. Two independent reviews conclude the models substantially overlap, with one arguing the field should "simplify models and so simplify treatments." And the decisive practical datum: "there is currently no evidence that any of the newer treatments are superior to standard CBT" — confirmed in the proponents' own RCT, between-treatment d = .07 at post-treatment and d = .12 at two years.

2.4 Attentional Control Theory: the one prediction that held

Eysenck et al. (2007) predicted that anxiety impairs the goal-directed attentional system — damaging inhibition and shifting — but that "anxiety may not impair performance effectiveness (quality of performance) when it leads to the use of compensatory strategies (e.g., enhanced effort)."

The definitive test (Shi, Sharpe & Abbott, 2019; PROSPERO-registered, 58 studies, N = 8,292, overall g = −0.58) confirmed it verbatim: anxiety produced "significant deficits in AC efficiency but not effectiveness; these deficits occurred in inhibition and switching but not updating; and studies with high cognitive load conditions found larger anxiety related AC deficits."

This is among the better-corroborated theoretical predictions in the anxiety literature — and it is routinely misused. ACT predicts preserved output at higher processing cost. It does not predict that anxious people perform worse. Applied literatures in education, occupational selection and clinical outcome almost universally invoke ACT to explain effectiveness decrements, which is the one thing the meta-analysis found anxiety does not reliably produce.

Note also that updating — the executive component most directly implicated in multi-step arithmetic and sustained academic work — was spared. That has direct consequences for §2.9.

2.5 Working memory

Moran (2016), 177 samples, N = 22,061:

Task familygkN
Overall−0.334 (p < 10⁻²⁹)17722,061
Complex span (e.g. OSPAN)−0.342303,196
Simple span (e.g. digit span)−0.31812717,547
Dynamic span (e.g. N-back)−0.437201,318

Two adversarial observations. Simple span carries 80% of the total N and shows the smallest effect — the pooled estimate is dominated by a task that makes minimal executive demands, which is awkward for an attentional-control account. And this is overwhelmingly self-reported anxiety correlated with task performance, so shared method variance and negative-affect-driven self-report bias are unexcluded. Moran himself frames the paper as identifying "methodological limitations common in the literature."

2.6 Learning under uncertainty: from landmark to null in ten years

The landmark. Browning et al. (2015) found that low-trait-anxious participants matched their learning rate to environmental volatility as a Bayesian model predicts, while high-trait-anxious individuals "showed less ability to adjust updating of outcome expectancies between stable and volatile environments," with reduced pupil-dilation sensitivity to volatility. It is a beautiful result and it launched a subfield.

The arc since:

StudyDesignResult
Gagne et al. (2020)Two experiments, bifactor modellingEffect present — but driven by a common internalizing factor, not anxiety-specific factors
McCoy & Lawson (2025)Original N = 80 + preregistered replication N = 160Volatility effects replicated only under low noise; anxiety effects "presented differently across experiments"
Satti et al. (2025) (preprint)Eight experiments, N = 820"No convincing evidence of a systematic relationship between internalizing symptoms and either learning rates or task performance… These findings challenge prominent claims that learning difficulties are a hallmark feature of internalizing psychopathology."

Gagne's result is the more interesting one to sit with: it does not refute the effect, it reassigns it — to transdiagnostic internalizing rather than anxiety. That is the same move the field is forced into repeatedly (see §2.3 on IU, §3.7 on structure).

Risk aversion. There is no meta-analysis of risk-taking in anxiety disorders. The available primary work is small and conditional: in N = 87, high and low socially anxious participants did not differ on any dependent variable absent a stressor — the group difference emerged only because low-anxious participants increased risk-taking under stress. Any claim that "anxiety causes risk aversion" as a main effect is unsupported by a well-powered synthesis.

Fear conditioning is the one paradigm with a robust and clinically resonant finding: across 44 studies, 963 patients and 1,222 controls, anxiety patients show elevated fear responses to conditioned safety cues (CS−) during acquisition and stronger responses to CS+ during extinction — i.e. impaired safety learning and delayed extinction. [unverified: the exact effect sizes and CIs are behind a paywall and no citing open-access source quotes them. Anyone citing a specific d from this paper should be asked to show the table.] The directions matter clinically because they are what exposure therapy exists to change.

2.7 Avoidance and safety behaviours

The claim that avoidance rather than arousal drives impairment is central to every cognitive-behavioural model of anxiety and is thinly quantified. I found no well-powered study directly pitting avoidance against arousal as the driver of functional impairment. It should be marked theoretically central but empirically under-tested.

On the more specific and heavily debated question — should safety behaviours always be eliminated during exposure? — the two best trials converge on no main effect either way:

  • N = 60 with clinical spider fear, exposure with elimination of safety behaviours vs exposure with judicious safety behaviours: "large effects on all measures from pretreatment to posttreatment… There were no significant group differences in treatment outcome or treatment acceptability."
  • N = 56, faded vs unfaded safety behaviours: approach and acceptability improved, but "this had little impact on treatment outcomes, with all groups showing similar benefits." Moderation: unfaded use benefited highly distress-intolerant participants; fading undermined exposure when it blocked expectancy violation.

Honest reading: clinical dogma ("always eliminate") and its challenger ("judicious use helps") are both under-supported. The answer is probably moderated by distress tolerance and by whether the behaviour blocks expectancy violation — and samples of 56–60 cannot resolve a moderator question.

2.8 Yerkes–Dodson does not survive contact with its source

Almost everyone who invokes the inverted-U has not read the 1908 paper. It takes twenty minutes and permanently changes how you cite it.

What it actually did. Yerkes & Dodson used Japanese dancing mice in a black/white brightness discrimination, punished by electric shock from a calibrated inductorium. Verified against the original: "Each of the forty mice experimented with…" Set III — the difficult discrimination, the only condition producing the famous curve — used "only one pair of dancers… with any given strength of stimulus," i.e. two mice per data point.

The easy condition was monotonic. Verbatim: "the rapidity of learning… increased as the strength of the stimulus increased. The weakest stimulus (135 units) gave the slowest rate… the strongest (420 units), the most rapid."

The paper contains zero inferential statistics and never uses the word "arousal." The authors themselves wrote: "Had we trained ten mice with each strength of stimulus instead of four the curve probably would have fallen regularly."

So the "law" is a difficulty × stimulus-intensity interaction in habit acquisition in mice, with n = 2 per cell in the critical condition. It is not a general arousal–performance function, and it was never about anxiety.

The drift has been documented, tracking disciplinary fashion through punishment, reward, motivation, drive, arousal, anxiety and stress. Two independent reviews call for its retirement, one noting that "the YDL has no basis in empirical fact but continues to inform managerial practices."

What modern evidence supports. There is reasonable evidence for an inverted-U between physiological arousal and simple perceptual-motor performance (quadratic trend accounting for 13.2% of variance, with a non-significant linear trend). There is essentially no good evidence for an inverted-U between self-reported anxiety and cognitive or academic performance: in N = 730 university students, "No evidence was found for a curvilinear relationship between arousal and performance."

Facilitative anxiety is confounded with self-confidence. In the flagship study of 48 gymnasts, "the only significant predictor of beam performance was self-confidence intensity" — the anxiety-direction variable did not survive its own regression. Meta-analytically (k = 48): cognitive anxiety r = −0.10, self-confidence r = +0.24. The construct is plausibly "positive emotions mislabelled as facilitative anxiety."

Choking is real and the mechanism is genuinely interesting: outcome-contingent pressure impairs rule-based learning while monitoring-by-others pressure impairs information-integration learning — a double dissociation between distraction and explicit-monitoring accounts. But the evidentiary base is thin: the foundational paper reports no effect sizes at all, with n = 18 per cell in the key experiment. No preregistered direct replication was located.

2.9 Test and math anxiety: reciprocal, small, and not mediated by working memory

The flagship test-anxiety meta-analysis publishes no numbers. Its abstract contains no r, no k-per-outcome, no N, no CI, no heterogeneity statistic. The only verifiable figure is "238 studies." Every specific effect size attributed to it is unverified.

Math anxiety, where the estimates exist, are not a constant:

Meta-analysisPooled rScope
Ma (1999)−.2726 studies
Namkung et al. (2019)−.34131 studies
Barroso et al. (2021)−.28 [−.29, −.26]747 effect sizes, I² = 90.4%, Egger p = .01
Caviola et al. (2022)−.30 [−.32, −.28]169 samples, N = 906,311
Finell et al. (2022)−0.168 [−0.203, −0.133]Egger p < .001; trim-and-fill → −0.136

That is a fourfold spread in variance explained across overlapping literatures.

Does working memory mediate? No. The largest test (15 independent samples) found the indirect effects "all non-significant and negligible in terms of magnitude (B ≤ .05 for all samples)." The paper universally cited as the mediation evidence describes no mediation analysis and reports only a "possible mechanism." And §2.4 supplies the theoretical reason: anxiety spares updating, the executive component arithmetic most needs.

Direction of causation is symmetric. A meta-analytic SEM of 67 samples and 39,935 participants: prior math anxiety → later math performance b = −.11; prior performance → later anxiety b = −.12. Verbatim: "MA is just as likely to be a consequence of MP as it is a cause."

(On stereotype threat, included because it is adjacent and frequently invoked: trim-and-fill reduces the girls-and-maths effect from g = −0.22 to g = −0.07, p = .27; a preregistered replication with N = 2,064 found P(H₀|x) = .963; and in real operational testing settings the estimate is d = −.01. [contested: p-curve analyses still find evidential value, and the race-based paradigm remains under-tested.])

2.10 Anxiety sensitivity: the best-replicated risk finding here

Amid the wreckage, one construct has held up.

  • Prospectively, in N = 1,401 Air Force trainees over five weeks, anxiety sensitivity predicted spontaneous panic after controlling for panic history and trait anxiety: "approximately 20% of those scoring in the upper decile on the ASI experienced a panic attack… compared with only 6% for the remainder" — a relative risk of roughly 3.3. Replicated in N = 1,296, and extended over two years in N = 404 to incident anxiety disorders in people with no history.
  • It survives the neuroticism control that killed intolerance of uncertainty: in the same latent-variable analysis where IU showed no unique association with any disorder, "anxiety sensitivity accounted for substantial unique variance" in depression, social anxiety, PTSD and panic.

Three qualifications that should travel with it. (i) It does not discriminate among anxiety disorders — anxiety patients vs controls d = 1.61, but panic vs PTSD d = 0.04. (ii) It relates more strongly to distress than to fear disorders, the opposite of what expectancy theory predicts. (iii) Moving it does not durably move outcomes: AS reduction is d = 0.54 post-treatment and d = 0.29, non-significant, at long-term follow-up.

So: a good marker, an unproven mechanism.

2.11 Real-world impact, and how much of it is comorbidity

Days out of role (24 countries, 62,971 interviews). The distinction between crude and adjusted figures is the whole story:

DisorderCrude days/yearAdjusted additional daysPopulation attributable risk
Panic disorder42.914.32.6%
PTSD42.715.22.2%
GAD39.87.71.0%
Social phobia39.37.31.7%
Specific phobia33.83.91.8%

The crude figure is roughly three times the adjusted one. Comorbidity, not anxiety alone, carries most of the burden — a point made starkly elsewhere: work-loss days per month per 100 workers were 49 for comorbid conditions versus 11 for a pure single disorder.

Earnings. In the US National Comorbidity Survey Replication, serious mental illness predicted $16,306 lower earnings — but "other 12-month and lifetime DSM-IV/CIDI mental disorders did not."

Educational and occupational attainment: much weaker than assumed. GAD, social phobia and PTSD were non-significant at all four educational milestones (12 of 12 p > .05). In low- and middle-income countries anxiety was protective for educational attainment (any anxiety OR 0.7). And in a twin design, persistent internalizing → unemployment fell from HR 1.56 (1.27–1.92) in the whole cohort to HR 0.96 (0.57–1.62), p = .88 in exposure-discordant twins — i.e. familial confounding accounts for it.

Healthcare utilisation is where anxiety's real-world footprint is largest and least discussed. Panic disorder is present in roughly 25–34% of emergency-department non-cardiac chest pain presentations, and in one study 98% of panic patients were unrecognised by attending cardiologists. And the intuitive remedy does not work: across 14 RCTs and 3,828 patients, diagnostic testing to reassure produced illness worry OR 0.87 (0.55–1.39), nonspecific anxiety SMD 0.06 (−0.16 to 0.28), symptom persistence OR 0.99 (0.85–1.15) — all null. Negative test results do not reassure anxious patients.

Driving. [contested, and the popular direction is wrong] Across 24 studies, there is "no category of disorder that was consistently associated with increased MVC risk." Every positive anxiety estimate uses a screening questionnaire plus self-reported crashes; the one study with objective diagnosis and coroner records gave OR 0.64 and 0.69, both non-significant. What does raise crash risk is the medication: benzodiazepines OR 1.59 (1.10–2.31), and with alcohol OR 7.69 (4.33–13.65).

(Contrast the stress review's finding that driving while observably emotional carried OR 9.8 in naturalistic data. Acute emotional state is a large per-episode risk; anxiety as a diagnosis is not.)

2.12 Anxiety versus stress and burnout: cost signature vs capacity signature

This is the most useful cross-review comparison available, and it is worth stating precisely because the two literatures reached it independently.

AnxietyAcute stressClinical burnout
Working memoryg = −0.334 (N = 22,061)g ≈ −0.20g = −0.36
InhibitionImpaired in efficiency onlyNot impaired overall (g = −0.076)Impaired
Shifting / switchingImpaired (efficiency)Impaired
UpdatingNot impaired
Memory retrievalNot establishedg ≈ −0.22g = −0.36
Broad fluid abilityNot establishedg = −0.36 to −0.53
Crystallised abilityNot establishedIntact
Accuracy / effectivenessPreserved via compensatory effortDegraded

The convergence. The stress review concluded — from bioenergetics, from the death of ego depletion, and from Wiehler's lateral-PFC glutamate finding — that stress raises the price of effort rather than removing the capacity for it. The anxiety literature, working from an entirely different tradition (Eysenck's attentional control theory, tested across 58 studies), reached the same structural conclusion: efficiency falls, effectiveness holds, and the mechanism is inferred to be compensatory effort.

Two literatures, different methods, different decades, same distinction. That convergence is the strongest reason to take the effort-cost framework seriously as a general account rather than a local finding.

The divergence, and it is the clinically useful part. Clinical burnout produces the opposite profile: a genuine ability-level decrement in fluid reasoning (−0.36 to −0.53) with crystallised knowledge spared — degraded output, not costlier output. Acute stress sits closer to anxiety in magnitude but differs in locus: it selectively hits retrieval and spares inhibition, whereas anxiety hits inhibition and shifting (in efficiency) and spares updating.

Do not over-read the table. The three literatures use non-comparable designs — anxiety findings are overwhelmingly correlational trait self-report × task performance; acute-stress findings are within-subject experimental; burnout findings are clinical group comparisons. Method variance alone could generate part of the divergence.

The comparison the data can support is narrow but real: there is no evidence that anxiety produces the broad fluid-ability decrement that characterises clinical burnout, and there is positive meta-analytic evidence that anxiety's cost falls on speed rather than accuracy.

Why this matters practically. If someone reports that work has become impossibly hard, the two profiles imply different things. Costlier-but-intact output means the intervention target is the cost structure — demand, control, recovery, competing load. Degraded output means the target is the impairment itself, and pushing harder makes it worse. The distinction is not diagnosable from the person's self-report, because subjective and objective impairment come apart in both conditions.

Part III — Taxonomy, epidemiology and course

3.1 What the manuals did, and the moving-target problem

DSM-5 (2013) removed OCD and PTSD from the anxiety disorders chapter. OCD went to a new Obsessive-Compulsive and Related Disorders chapter; PTSD and acute stress disorder went to Trauma- and Stressor-Related Disorders. Two disorders moved in — selective mutism and separation anxiety disorder. Panic disorder and agoraphobia were uncoupled into separately codable diagnoses.

The rationales were: for OCD, that compulsivity rather than anxiety is the core feature; for PTSD, that factor analyses supported four rather than three symptom clusters, with dysphoric, aggressive, guilt-shame and dissociative symptoms beyond the fear-based ones.

The criticism is sharp, and it includes the workgroup's own. A pre-emptive review concluded reclassifying OCD was "premature and not supported by the currently available data." A post-hoc review went further: the OCRD class's "empirical validity and practical utility are questionable," leading to "a rejection of the OCRD classification on both scientific and logical grounds." Most tellingly, the DSM-5 sub-workgroup's own published preliminary recommendation was to keep OCD inside the anxiety disorders category. The final manual did not follow it.

ICD-11 (in effect since January 2022) creates "Anxiety and fear-related disorders," organised by the focus of apprehension — the stimulus the person reports as triggering anxiety — replacing ICD-10's phobic/other split. Three differences from DSM-5 matter:

  • Hierarchical exclusion rules were abolished. GAD may now be diagnosed alongside depressive disorders, phobias and OCD. This alone will raise measured GAD prevalence and comorbidity relative to ICD-10.
  • Mixed depressive and anxiety disorder moved out of anxiety into depressive disorders — the reverse direction to DSM's moves.
  • Agoraphobia is defined more broadly (fear of incapacitating or embarrassing outcomes, not fear of open spaces).

The moving-target problem — the most consequential point in this section

Every prevalence estimate published before roughly 2013 that reports "anxiety disorders" includes OCD and PTSD. The 2018 WHO treatment-gap paper still uses DSM-IV and therefore still counts PTSD as an anxiety disorder. The canonical sex-ratio analysis includes PTSD — and PTSD has the largest female excess of any condition in that dataset (OR 2.6), so its inclusion inflates the headline anxiety sex ratio. GBD models anxiety as a bundle whose composition has shifted across rounds.

Any trend claim spanning 2013 is comparing partly different constructs, and almost no published trend analysis adjusts for this. Bear it in mind through §3.5.

3.2 The GAD threshold: a policy decision dressed as nosology

DSM-5-TR retains the 6-month duration requirement for GAD. (Flag: at least one peer-reviewed source states the criterion "decreased to 3 months in DSM-5." That is incorrect — it describes a pre-publication draft proposal that was rejected. Reviewers repeating it are propagating an error.)

The evidence against six months is unusually clean. Using WHO data across 17 countries (N = 85,052), lifetime GAD prevalence at 1/3/6/12-month thresholds was 7.5% / 5.2% / 4.1% / 3.0% in developed countries — with "little difference between GAD of 6 months' duration and GAD of shorter durations… in age of onset, symptom severity or persistence, co-morbidity or impairment." Only the ≥12-month subgroup was distinctively severe. Independently, in the prospective Zurich Cohort Study the 6-month criterion "could not be confirmed as clinically meaningful" and would preclude diagnosis in about half of subjects actually being treated for generalized anxiety syndromes.

A proposal to relabel GAD "generalized worry disorder" with a 3-month threshold was not adopted.

Which way does the threshold err? Both. The duration criterion excludes a large, equally impaired group. Simultaneously the DSM-5 definition is more inclusive than DSM-IV. And underneath: worry has a dimensional, not taxonic, latent structure — meaning there is no natural cut point to discover, only one to choose. Where the line falls is a policy decision presented as a nosological finding.

3.3 Prevalence, done carefully

Be explicit about type. These are not interchangeable and are routinely conflated:

EstimateTypeValueSource
Global, methodology-adjustedPoint/current7.3% (4.8–10.9)Baxter et al., 2013
Global, GBD modelPoint3.78% (3,779.5/100,000), 2019GBD 2019
Cross-national, 21 countries12-month9.8% (range 3.0% to 19.0%)Alonso et al., 2018
Cross-national, 17 countriesLifetime4.8–31.0% (IQR 9.9–16.7%)Kessler et al., 2007
DSM-5 GAD, 26 countriesLifetime / 12-mo / 30-day3.7% / 1.8% / 0.8%Ruscio et al., 2017

Note the twofold discrepancy between Baxter's adjusted current prevalence (7.3%) and GBD's modelled point prevalence (3.78%) — from overlapping author groups using overlapping data. The gap is methodological, and reviewers quote whichever suits.

The high-income gradient is large and consistent — current prevalence 5.3% in African cultures vs 10.4% in Euro/Anglo cultures; lifetime GAD 5.0% high-income, 2.8% middle, 1.6% low. Three explanations compete and none is decisive:

  1. True difference. But note the paradox that undermines a simple "affluence causes anxiety" story: GAD is negatively associated with SES within countries while being positively associated with national income between countries.
  2. Measurement and reporting artefact. Methodological factors explained an additional 13% of between-study variance after substantive factors, and the authors concluded that "specific attention should be paid to cultural differences in responses to survey instruments for anxiety disorders." Willingness to disclose internal distress to a lay interviewer is not culturally invariant.
  3. Concept translation. Diagnostic interview items were developed in English in high-income settings. ICD-11 explicitly incorporated cultural concepts of distress (ataque de nervios, khyâl cap, trúng gió) because standard panic criteria do not map cleanly onto how the experience is organised elsewhere.

Verdict: the gradient is real in the data and unresolved in interpretation. No published study cleanly partitions it. It should not be presented as evidence of a true aetiological difference.

3.4 The COVID question: modelled versus observed

The headline. The COVID-19 Mental Disorders Collaborators (2021) estimated "an additional 76·2 million (64·3 to 90·6) cases of anxiety disorders globally (an increase of 25·6% [23·2 to 28·0])."

What that estimate rests on. From 5,683 records screened, 27 studies met inclusion for anxiety. These were fed into a meta-regression linking prevalence change to two ecological proxies — human mobility and daily infection rate — and extrapolated to all 204 countries and territories, the overwhelming majority of which contributed no primary data at all. The narrow uncertainty interval reflects parameter uncertainty inside the model, not uncertainty about whether the model's structure is right.

The counter-evidence is a design argument, and it is strong. Restrict to studies measuring the same people before and during:

  • 65 longitudinal cohorts: symptom change March–April 2020 SMC 0.102 (0.026–0.192); by May–July 2020 0.067 (−0.022 to 0.157) — no longer significant.
  • 137 studies / 134 cohorts with ≥90% retention: general population anxiety symptoms SMD 0.05 (−0.04 to 0.13) — not statistically significant. (Depression 0.12, significant; women's anxiety 0.20, significant.)

Population-level anxiety did not detectably move once you follow the same people. That is the single most important finding in this section.

What happened after. UK data show anxiety declining across the first 20 weeks of lockdown (b = −1.93, p < 0.0001), fastest in weeks 2–5; the association with policy stringency shrinking across phases; and 76.8% of people showing consistently good mental health April–October 2020, with ~11% sustained or worsening. State-dependent and largely reverting, not a permanently shifted baseline.

But the estimate has been baked into the official series. WHO's fact sheet now states "In 2021, 359 million people were living with an anxiety disorder," citing GBD 2021. GBD 2019's published figure was 301.4 million. A ~19% jump in two years is not new primary epidemiology; it is the COVID adjustment propagating forward. Anyone citing 359 million alongside the null same-cohort finding is citing two mutually inconsistent claims.

3.5 Is there an anxiety epidemic? Three claims, kept separate

Almost all public discussion of this question fails by conflating three different propositions. Keep them apart:

(a) Rising diagnosis and help-seeking — enormous, and largely uninformative about (c)

UK primary care, N = 9,133,246, ages 1–20, 2003–2018: anxiety disorder diagnosis incidence rate ratio 3.51 (3.18–3.89). First-ever antidepressant prescriptions in UK 3–17-year-olds nearly doubled 2006–2015, with only 21% of 2015 incident prescriptions linkable to a depression diagnosis.

And the decisive demonstration of what this metric measures: in April 2020, UK anxiety diagnosis incidence fell 47.8% (44.3–51.2) below expected in a single month. Nobody thinks anxiety halved. Diagnosis rates track service access.

(b) Rising self-reported symptoms — real but modest, and instrument-dependent

Eight of eleven studies using the General Health Questionnaire found significant increases in psychological distress over time. This is from the very paper usually cited to debunk the epidemic — Baxter et al. (2014) did not find "nothing changed." They found diagnosed disorder prevalence flat while symptom-scale distress rose. That is the prevalence inflation hypothesis, stated five years early by the people cited against it.

(Note that most of the widely publicised youth trend data — Twenge and successors — concerns depression, not anxiety.)

(c) Rising true disorder prevalence — one genuinely clean study, and it is positive

Taxiarchi et al. (2026) is the study that matters: same country, same DAWBA diagnostic interview, probability samples, N = 13,561, ages 5–16, 2004 versus 2017. Anxiety disorders 3.5% → 5.4%, RR 1.63 (1.37–1.93).

Two caveats the authors supply. The rise was concentrated in White children (RR 1.88) with no clear change in minority-ethnic children (RR 0.85, 0.52–1.39) — a pattern hard to explain by uniform true-incidence increase, and suggestive of differential recognition or reporting. And it is one country, one instrument, a thirteen-year gap.

Supporting evidence cuts both ways: SDQ impact scores rose 1999–2017 specifically among already-diagnosed children, which cuts against a pure "we're catching milder cases" story. Against that, other analyses find secular change including "periods of increase and decrease" — non-monotonic.

The baseline against which all this should be read: age-standardised anxiety point prevalence was 3.8% (1990) → 4.0% (2010), with crude case counts rising 36%, fully explained by population growth and ageing.

Prevalence inflation

Foulkes & Andrews (2023) propose that awareness efforts may themselves raise reported prevalence via two mechanisms: overinterpretation — people "interpret and report milder forms of distress as mental health problems" — and a self-fulfilling prophecy, where "labelling distress as a mental health problem can affect an individual's self-concept and behaviour in a way that is ultimately self-fulfilling." Their anxiety-specific illustration is direct: "interpreting low levels of anxiety as symptomatic of an anxiety disorder might lead to behavioural avoidance, which can further exacerbate anxiety symptoms."

That mechanism is theoretically well-motivated — avoidance is the best-supported maintaining factor in every cognitive-behavioural model. But three years on, no published study directly testing the hypothesis with new data was located. It remains a call to test that has not been answered.

Reading. Claim (a) is enormous and mostly measures services. Claim (b) is real and modest. Claim (c) has exactly one methodologically clean positive result — genuinely important, but single-country with an ethnicity pattern that itself hints at recognition effects. "Epidemic" is not supportable. "A real but modest rise in youth anxiety in some high-income countries, embedded in a much larger rise in labelling" is.

3.6 Age of onset: true for the block, false for GAD

Across 192 studies and N = 708,561, the anxiety/fear-related block has the earliest onset of any disorder class: 38.1% before age 14, 51.8% before 18, 73.3% before 25; peak age 5.5 years (median 17).

The adversarial point that is almost always omitted: this is true for the block and false for several of its members. Disorder-specific onsets place phobias, separation anxiety and social anxiety at 8–13 years, but panic disorder at 25–27 and generalized anxiety disorder at 30–35. GAD clusters with post-traumatic, depressive and bipolar disorders, not with the phobias. The block's peak of 5.5 years is driven by specific phobia and separation anxiety.

Clinical-cohort data agree: mean age at illness onset was 14.4 years for social phobia but 21.3 for GAD, 26.8 for panic with agoraphobia and 34.0 for panic without agoraphobia.

Implication, and it is directly actionable: early-intervention programmes targeting "anxiety" in childhood are targeting phobias and separation anxiety. They will not prevent GAD, whose modal onset is two decades later. Any prevention strategy justified by the "anxiety starts at 5" statistic is aimed at the wrong disorders.

3.7 Structure: where GAD actually belongs

Comorbidity is the norm, not the exception. Of those with a current anxiety disorder, 63% had current and 81% lifetime comorbid depression. For GAD specifically, lifetime comorbidity is 81.9%. Anxiety preceded depression in 57% of comorbid cases, depression preceded anxiety in 18%.

The genetics are close to unity. SNP-based genetic correlations with depression: rg = 0.78, 0.81, 0.9, and in the largest GWAS to date rg = 0.91. Twin estimates are higher still: rg = +1.00 in females and +0.74 in males between major depression and GAD, with neuroticism accounting for only ~25% of the shared genetic covariance.

The structural models place GAD with depression. Clark & Watson's tripartite model posited physiological hyperarousal as the anxiety-specific component — but the hyperarousal component does not generalise: in 350 outpatients, the GAD latent factor related to autonomic suppression, not hyperarousal. The internalising spectrum splits into distress (major depression, dysthymia, GAD, PTSD) and fear (panic, agoraphobia, social phobia, specific phobia) subfactors.

This is the deepest problem for the DSM/ICD anxiety category. Quantitative structural evidence places GAD closer to major depression than to panic or the phobias. The manuals put GAD with panic and the phobias. Both cannot be right about carving nature at its joints.

The p factor and its critics. [contested] A single general dimension fits better than three in several datasets. But bifactor models may out-fit alternatives as an artefact of model flexibility, and the most damaging finding is that fear/distress specific factors showed poor reliability and validity, while a raw count of comorbid diagnoses performed comparably to p on most validity tests.

Are these natural kinds? Taxometric evidence is not univocal. Worry is dimensional. Social anxiety is contested — dimensional in one analysis, taxonic in another. No consensus exists.

Put §3.7 together with Part I and a coherent, uncomfortable picture emerges. The failure to find a pathophysiology of anxiety (Part I) and the failure of the category to survive quantitative structural analysis (here) may be the same fact seen twice. You will not find the biology of a category that is not a biological kind.

3.8 Course, recurrence, and why published relapse rates are too low

Twelve-year follow-up of a clinical cohort (N = 711 at intake, 473 at year 12), Kaplan-Meier cumulative probabilities:

DisorderP(recovery) by year 12P(recurrence) after recovery% of follow-up in episode
Panic without agoraphobia0.820.5641%
Generalized anxiety disorder0.580.4574%
Panic with agoraphobia0.480.5878%
Social phobia0.370.3980%

Comorbid major depression roughly halved the likelihood of recovery from panic-with-agoraphobia and GAD, and nearly doubled recurrence risk.

Do not over-generalise this. It is a treatment-seeking sample from ~30 clinician practices, 3% minority, DSM-III-R criteria, 33% attrition. The very low recovery rates reflect chronicity-enriched clinical recruitment, not the community course.

Community/primary-care course is better: pure anxiety median episode duration 16 months, 41.9% chronic at 2 years — worse than pure depression (6 months, 24.5% chronic) and better than comorbid anxiety-depression (>24 months, 56.8% chronic).

And published recurrence rates are underestimates by construction. Counting only the index disorder gives 2-year recurrence of 23.5%. Over four years, anxiety recurrence rose from 23.8% (diagnostically stable) to 54.8% once newly arisen anxiety or depressive disorders were counted. Most people who "relapse" do so into a different diagnosis, and most studies do not count that.

Anxiety → later depression: the directional claim is overstated. Across 66 studies and N = 88,336: anxiety symptoms → later depressive symptoms r = .34; depressive → later anxiety r = .31. At disorder level, anxiety → depression OR 2.77; depression → anxiety OR 2.73. The asymmetry was "very small and likely not clinically meaningful." Two exceptions run the other way: depression → social anxiety OR 6.05 and → specific phobia OR 2.93.

The common framing of anxiety as "a gateway to depression" is not supported. The relationship is essentially symmetrical — which is what you would expect given rg = 0.91.

3.9 Physical health outcomes

OutcomeEstimateStudies / N
Incident coronary heart diseaseHR 1.26 (1.15–1.38)20 / 249,846, 11.2y
Cardiac deathHR 1.48 (1.14–1.92)same
Nonfatal MIHR 1.43 (0.85–2.40) — nssame
Incident cardiovascular diseaseHR 1.52 (1.36–1.71)37 / 1,565,699
Incident strokeHR 1.24 (1.09–1.41)8 / 950,759
Incident hypertensionadj. HR 1.55 (1.24–1.94)8 / 80,146
IBS onsetRR 2.38 (1.58–3.60)11

In established cardiac disease, adjustment dissolves most of it. Across 44 articles and N = 30,527, "when adjusting for covariates, nearly all associations became nonsignificant," with effects present in stable coronary disease but not post-acute-coronary-syndrome. One post-PCI cohort even found anxiety protective — a single unreplicated result, but a caution against directional certainty.

All-cause mortality: the association largely dissolves. Overall HR 1.09 (1.01–1.16) — but publication-bias-adjusted HR 1.03 (0.95–1.13), ns; community samples only 0.99 (0.96–1.02); studies adjusting for depression 1.01 (0.96–1.06). Against this, Danish nationwide register data on clinically diagnosed anxiety (>30 million person-years) found natural-cause mortality rate ratio 1.39 and unnatural-cause 2.46, rising to 11.72 with comorbid depression.

Reconciliation: severe, clinically diagnosed anxiety with comorbid depression carries excess mortality, largely from unnatural causes. Community-level anxiety symptoms, adjusted for depression, do not.

Dementia is genuinely unresolved, and every major meta-analysis says so in its own conclusions. The community signal (RR 1.57) was "driven by studies with mean age of 80 years or above," which the authors read as prodrome. The direct test of reverse causation — requiring ≥10 years between anxiety assessment and dementia diagnosis — found only four eligible studies from 3,510 screened, with mixed results.

3.10 Treatment gap, and an audit of the cost figures

The gap is large, and mostly demand-side. Across 21 countries (N = 51,547):

OverallHigh-incomeLow / lower-middle income
Perceived need for treatment41.3%48%28.5%
Received any treatment27.6%36.3%13.1%
Received "possibly adequate" treatment9.8%13.8%2.3%
Treated, among those perceiving need66.8%75.0%46.1%

Note the structure: fewer than half of people with a 12-month anxiety disorder perceive a need for care, but two-thirds of those who do get some treatment. The dominant bottleneck is recognition and perceived need, not access alone. For comparison, minimally adequate treatment for major depression in the same infrastructure was 16.5% — anxiety is treated less adequately than depression.

Cost figures — the provenance audit, and it is damning:

FigureActual sourcePeer-reviewed?Verdict
"US$1 trillion/year, depression + anxiety"WHO press release, 2016NoThe number does not appear in its usual citation. That paper reports 15-year (2016–30) net present values: $147bn investment, $230bn depression + $169bn anxiety productivity gain, $310bn health value. No arithmetic path leads to $1tn/year.
"$6 trillion by 2030"WEF/Harvard report, 2011No — grey literatureCovers all mental illness, not anxiety+depression
"~US$5 trillion (2019)"Arias, Saxena & Verguet, 2022YesValue-of-Statistical-Life method; 418M DALYs, which the authors call "more than three-fold" the GBD estimate. Explicitly maximalist

All three use different methods, scopes and time bases. Do not quote any of them without stating which.

Burden. Mental disorders are now the leading global cause of years lived with disability (17.3%), with anxiety disorders ranking 11th among 304 causes for DALYs and peaking in the 15–19 age group. (Note a methodological instability worth flagging: GBD 2019 reported age-standardised anxiety prevalence as flat; GBD 2023 reports "notable increases" in the same metric. That is a finding about the models, not only about the world.)

Part IV — What actually helps

4.1 The comparator rule

Before any table: the cleanest demonstration in the field of why comparator choice dominates everything. Social anxiety disorder, 101 trials, 13,164 participants:

Interventionvs WAITLISTvs APPROPRIATE PLACEBO
Individual CBT−1.19 (−1.56 to −0.81)−0.56 (−1.00 to −0.11) vs psychological placebo
SSRIs / SNRIs−0.91 (−1.23 to −0.60)−0.44 (−0.67 to −0.22) vs pill placebo
MAOIs−1.01 (−1.56 to −0.45)not established vs placebo
Benzodiazepines−0.96 (−1.56 to −0.36)not established vs placebo
Group CBT−0.92 (−1.33 to −0.51)not established vs placebo
Self-help with support−0.86 (−1.36 to −0.36)not established vs placebo
Self-help without support−0.75 (−1.25 to −0.26)not established vs placebo
Psychodynamic therapy−0.62 (−0.93 to −0.31)not established vs placebo

Verbatim: individual CBT and SSRIs/SNRIs "were the only classes of interventions that had greater effects on outcomes than appropriate placebo." Everything else beat a waitlist. Roughly half of each waitlist effect size disappears when a credible control is used.

Keep this table in mind for the rest of Part IV. It is the reason the tier assignments below are more conservative than most guidelines.

4.2 Evidence tier table

TierInterventionBest estimateComparator
STRONGIndividual CBT, social anxietySMD −0.56 (−1.00 to −0.11)Psychological placebo
STRONGSSRIs/SNRIs, social anxietySMD −0.44 (−0.67 to −0.22)Pill placebo
STRONGCBT, anxiety-related disorders pooledg = 0.56Psychological or pill placebo
STRONGAntidepressants, panic disorderRR 0.72 (0.66–0.79), NNTB 7 (6–9)Pill placebo (low-quality evidence)
STRONGDuloxetine / pregabalin / venlafaxine / escitalopram, GADHAM-A −3.13 / −2.79 / −2.69 / −2.45Pill placebo
STRONGAntidepressant continuationRelapse OR 3.11 (2.48–3.89) on discontinuationContinued drug vs placebo switch
STRONG (equivalence)Guided internet CBTg = 0.02 (−0.09 to 0.14)Face-to-face CBT
MODERATECBT, GADSMD −0.74 (−1.09 to −0.38)Treatment as usual
MODERATEMBSR, anxiety disordersDifference −0.07 (−0.38 to 0.23)Escitalopram 10–20 mg
MODERATEMindfulness meditationES 0.38 (0.12–0.64) at 8 weeksActive controls
MODERATEBenzodiazepines, panic / social anxietyTop-ranked for response; SMD −0.96Placebo (low quality) / waitlist
MODERATEQuetiapine, GADHAM-A −3.60 (largest of any agent); discontinuation OR 1.44Pill placebo
MODERATEKavaHAM-A WMD 5.0 (1.1–8.8)Placebo
MODERATESilexan (lavender oil)HAM-A 14.1 vs 9.5 placeboPlacebo (sponsor-conflicted; see §4.9)
MODERATEExercise, anxiety symptomsmedian ES −0.42 (IQR −0.66 to −0.26)Usual care
MODERATE (inactive only)Group ACTg = 0.52 (0.30–0.73)Non-active controls only — null vs CBT
LIMITEDApplied relaxation, GAD−0.59 (−1.07 to −0.11) → −0.47 (−1.18 to 0.23) low-RoB onlyTreatment as usual
LIMITEDExercise, diagnosed anxiety disordersSMD −0.582, 6 RCTs, n = 262, CI misprintedUsual treatment
LIMITEDApps for anxietyg = 0.26, NNT 12.4Mostly passive controls
LIMITEDInterpretation bias modificationSMD −0.30 (−0.50 to −0.10)Sham
LIMITEDPsilocybin, end-of-life anxietyn = 29 and n = 51; blinding compromisedNiacin / very low dose
NULLUnguided self-help appsg = 0.10 (ns) for anxietyActive / placebo apps
NULLAttention Bias Modificationg = 0.29 (−0.14 to 0.74) clinical; registered replication nullSham
NULLD-cycloserine augmentationd = −0.25 post-tx; d = −0.19, CI crosses 0 at follow-upPlacebo + exposure
NULLPropranolol, any anxiety disorder8 trials total; largest social-phobia study n = 16Placebo / benzodiazepine
NULLCannabinoids for anxietyNo significant effect; AEs OR 1.75, NNTH 7Placebo
INSUFFICIENTAshwagandhaSMD −1.55 to −6.87, I² 93–98% — literature not trustworthyPlacebo
INSUFFICIENTMagnesium / chamomile / weighted blankets / L-theanineNo adequate meta-analytic estimate in anxiety
INSUFFICIENTKetamine, anxiety disordersNo dedicated meta-analysis exists

4.3 CBT and exposure

Against placebo, CBT's effect is moderate, not large. Restricted to 41 randomised placebo-controlled trials (N = 2,843): target disorder symptoms g = 0.56; other anxiety symptoms 0.38; depression 0.31; quality of life 0.30; response OR = 2.97.

By disorder, large effects for OCD, GAD and acute stress disorder; small-to-moderate for PTSD, social anxiety and panic. Exposure-based protocols outperformed cognitive-only, but not significantly so. And note a real harm signal: PTSD dropout 29.0% CBT vs 17.2% placebo — CBT is not a low-harm intervention in trauma.

Long-term outcomes are thinner than the field admits. Across 69 RCTs and 4,118 outpatients, described by the authors as "mainly of low quality": gains hold reasonably through 6–12 months. At ≥12 months the evidence largely evaporates — GAD g = 0.22 (k = 10), social anxiety 0.42 (k = 3), PTSD 0.84 (k = 5); not significant for panic disorder (k = 5); not calculable for specific phobia (k = 1) or OCD (k = 0).

Relapse rates of 0–14% at 3–12 months were reported in only 6 of 69 trials. Selective reporting is near-certain; treat the low figure with suspicion.

The honest statement: we do not know whether CBT gains hold beyond a year for most anxiety disorders, because almost nobody has looked.

Exposure and the inhibitory-learning shift. Craske et al. (2014) reframed exposure from habituation (wait for fear to subside within session) to inhibitory learning (build a competing non-threat association; within-session fear reduction is not the target), with eight operational strategies including expectancy violation, deepened extinction, removal of safety signals, variability and multiple contexts.

The model is theoretically better grounded and fits the fear-conditioning data in §2.6. But the comparative claim is unverified: most support comes from analogue and laboratory extinction work, and no adequately powered clinical RCT establishing superiority of an integrated inhibitory-learning protocol over standard exposure was located.

Is exposure under-used? The defensible version: exposure is delivered, but often in diluted form, and the disorder-specific techniques with the strongest evidence are the ones most often omitted — interoceptive exposure for panic, self-directed in vivo exposure, exposure-and-response-prevention for OCD.

D-cycloserine augmentation: promise to near-null. Individual participant data from 21 of 22 eligible trials (1,047 participants): post-treatment d = −0.25 (p = .01); mid-treatment non-significant; follow-up d = −0.19, CI −5.99 to 0.03, i.e. crossing zero. No moderators found. A small, fragile augmentation effect at best. Not clinically deployable.

4.4 Pharmacotherapy

GAD — 89 trials, 25,441 patients, 22 drugs vs placebo:

DrugHAM-A mean difference (95% CrI)Tolerability
Quetiapine−3.60 (−4.83 to −2.39)Poor — discontinuation OR 1.44
Duloxetine−3.13 (−4.13 to −2.13)Good
Pregabalin−2.79 (−3.69 to −1.91)Good
Venlafaxine−2.69 (−3.50 to −1.89)Good
Escitalopram−2.45 (−3.27 to −1.63)Good
Paroxetine, benzodiazepinesEffectivePoorly tolerated vs placebo

Context for magnitude: roughly 2.5–3.6 HAM-A points on a scale where trial entry typically requires ≥18. A real but modest drug–placebo separation.

Panic disorder. Antidepressants vs placebo: failure to respond RR 0.72 (0.66–0.79), NNTB 7 (6–9) — low-quality evidence, 6,500 participants. In the Cochrane network, diazepam, alprazolam and clonazepam ranked highest for efficacy, and benzodiazepines had the lowest dropout — but at low evidence quality, and no class differed from another.

Discontinuation symptoms, done properly. An earlier widely publicised estimate (withdrawal incidence 27–86%, weighted average 56%) was heavily criticised for pooling self-selected online withdrawal-forum samples with clinical trials. The careful successor (79 studies, 21,002 patients, and critically including a placebo-discontinuation comparator):

  • ≥1 discontinuation symptom after antidepressant: 31% (27–35%)
  • ≥1 after placebo: 17% (14–21%)
  • Difference within RCTs: 8 percentage points (4–12)
  • Severe symptoms: 2.8% (1.4–5.7%) vs 0.6% (0.2–1.3%)
  • Authors' bottom line: ~15% net incidence — roughly 1 in 6 to 7 patients

Worst offenders by frequency: desvenlafaxine, venlafaxine, imipramine, escitalopram.

Relapse on discontinuation. Across 28 relapse-prevention RCTs (n = 5,233, low risk of bias): discontinuation OR of relapse 3.11 (2.48–3.89); time to relapse HR 3.63. Pooled relapse 36.4% placebo vs 16.4% continued antidepressant at up to one year. No moderation by disorder, drug, taper mode or treatment duration. Beyond one year, no evidence exists.

4.5 Benzodiazepines: the honest position

The efficacy is real and it is fast. In the panic network they ranked highest for response and were the only class with lower dropout than placebo. In social anxiety they beat waitlist at SMD −0.96. Both estimates come from low-quality evidence, and neither established superiority over an appropriate placebo.

Against that: chronic use is prevalent (~4% of the general population), discontinuation after long-term use is genuinely difficult, and no pharmacological taper aid has convincing support (38 trials, 2,543 participants, high risk of bias in all but one). What does work is psychological: CBT plus gradual tapering versus tapering alone gives NNT 3.2 at 3 months and NNT 2.8 at 6–12 months — though from only 3 RCTs.

Guideline position is unambiguous: "Do not offer a benzodiazepine for the treatment of GAD in primary or secondary care except as a short-term measure during crises."

Has the pendulum swung too far? An honest reading: the efficacy evidence is at least as strong as for SSRIs in panic and social anxiety, and the short-term tolerability evidence is better. What justifies the restriction is not weak efficacy but (a) dependence and discontinuation difficulty, (b) the near-total absence of long-term controlled data, and (c) cognitive, fall and motor-vehicle harms — recall from §2.11 that benzodiazepines, not anxiety, are what raises crash risk (OR 1.59; with alcohol 7.69).

[unverified] Quantitative dependence-risk estimates that would let a clinician give a patient a number are not available from high-quality prospective cohorts. This is a genuine evidence gap, and confident blanket statements in either direction are unsupported.

4.6 Propranolol: the clearest case of belief outrunning evidence

Beta-blockers for performance anxiety are near-universal among musicians, public speakers and surgeons. The evidence base is two small studies and a systematic review that says not to.

The origin. A 1982 paper reporting two trials, 29 subjects total, in musicians, with performance quality rated by critics. That is the primary evidence for a now-ubiquitous practice.

The systematic review. Eight studies total across all anxiety disorders: panic (4 studies, n = 130), specific phobia (2, n = 37), social phobia (1 study, n = 16), PTSD (1, n = 19). No significant difference versus benzodiazepines in panic. Conclusion, verbatim: "the quality of evidence for the efficacy of propranolol at present is insufficient to support the routine use of propranolol in the treatment of any of the anxiety disorders."

Verdict: the belief vastly outruns the evidence. This is the clearest example in the review of a practice sustained by plausible mechanism plus vivid anecdote. It may still work — a single 16-person trial cannot rule that out either — but nobody currently knows, and it is presented as settled.

Other agents, briefly. Pregabalin is genuinely supported for GAD (though misuse liability is not captured in the effect estimate). Buspirone is efficacious but "limited by small sample sizes." Quetiapine has the largest efficacy signal and significantly worse acceptability — guidelines say do not use it in primary care. Whether combination psychotherapy + pharmacotherapy beats monotherapy in anxiety disorders is not settled by high-quality evidence — Cochrane found inadequate evidence and no superseding synthesis exists.

4.7 Digital and low-intensity

Guided internet CBT is equivalent to face-to-face CBT. Across 31 RCTs in 16 conditions, pooled g = 0.02 (−0.09 to 0.14). This is an equivalence finding of real practical importance and it is the strongest argument for digital delivery in this entire review.

Guidance is the active ingredient, and it shows up in adherence. Long-term outcomes are equivalent (g = −0.07), but completion is 86.7% face-to-face, 79.1% therapist-guided iCBT, 48.2% self-guided iCBT. The social-anxiety network shows the same ordering against waitlist: self-help with support −0.86, without support −0.75 — a modest efficacy gap and a very large engagement gap.

Apps are a different proposition entirely. Across 176 RCTs: generalized anxiety g = 0.26 (N = 22,394), NNT 12.4. Some subgroups look better (social anxiety 0.52, acrophobia 0.90) and one is nominally negative (panic g = −0.12), all from small high-risk-of-bias trials.

The number the marketing omits: against active or placebo apps, mindfulness apps produce g = −0.15 for depression and g = 0.10 for anxiety — both non-significant. Against passive controls the same apps give anxiety g = 0.28. The entire app effect in anxiety may be attention, expectancy and measurement reactivity.

(This converges exactly with the stress review, which found app-for-stress g = 0.09 against a placebo app and real-world 30-day retention of 3.3%.)

4.8 Exercise, mindfulness, relaxation and breathing

Exercise. The only meta-analysis restricted to diagnosed anxiety and stress-related disorders comprises 6 RCTs and 262 adults, SMD = −0.582. ⚠️ The published confidence interval is arithmetically impossible — the point estimate lies outside its own stated interval — an uncorrected typographical error meaning the result is marginal and compatible with a near-null effect. Anyone citing "a moderate effect of exercise in diagnosed anxiety disorders" is citing a 262-person evidence base with a printing error in its CI.

For anxiety symptoms in mixed populations the umbrella estimate is median ES −0.42 — but 77 of 97 included reviews scored "critically low" on AMSTAR-2, and the comparator was usual care, not an attention-matched control. Exercise trials cannot blind.

Verdict: MODERATE for symptoms, LIMITED for diagnosed disorders. Still worth doing — the harm profile is excellent and the stress review found strong evidence for mood — but the anxiety-specific claim is weaker than commonly stated.

Mindfulness. The anchor deliberately included only trials with active controls: anxiety ES 0.38 (0.12–0.64) at 8 weeks, 0.22 at 3–6 months. And the sentence that should accompany every citation: "We found no evidence that meditation programs were better than any active treatment (ie, drugs, exercise, and other behavioral therapies)."

MBSR versus escitalopram is the strongest single result: in a 276-participant noninferiority RCT, CGI-S reduction 1.35 vs 1.43, difference −0.07 (−0.38 to 0.23), within the prespecified margin. Adverse events: 78.6% escitalopram vs 15.4% MBSR; 10 escitalopram dropouts for adverse events versus zero.

Two adversarial caveats. The noninferiority margin of 0.495 CGI-S points is wide — this establishes "not much worse," not "as good." And the subsequent videoconference phase failed to replicate: MBSR-VC vs escitalopram-VC, p = 0.17, with noninferiority NOT supported and escitalopram showing higher satisfaction and greater impact on panic symptoms.

ACT. Group ACT gives anxiety g = 0.52 — but the subgroup analysis is the point: ACT was superior to non-active controls for both anxiety and depression, but superior to active controls (e.g. CBT) only for depressive symptoms — not for anxiety.

Applied relaxation retains formal guideline status for GAD but the evidence has thinned: SMD −0.59 (−1.07 to −0.11) at low certainty, losing statistical significance when high-risk-of-bias trials are excluded (−0.47, −1.18 to 0.23). It was also the only Step-3 psychological option that did not survive to 3–12 month follow-up. Its guideline status is stronger than its current evidence.

The "physiological sigh." This deserves precision because it is heavily cited in popular media. The trial was a remote, randomised study of three 5-minute daily breathwork exercises versus an equal period of mindfulness meditation over one month. Finding: breathwork, especially exhale-focused cyclic sighing, produced greater improvement in mood (p < 0.05) and reduction in respiratory rate (p < 0.05) than mindfulness meditation.

Why the popular version is wrong: (a) the comparator is meditation, not placebo or no treatment — the claim is "slightly better than meditation," not "reduces anxiety"; (b) the headline positive outcomes are mood and respiratory rate; (c) participants were remote, unsupervised, self-selected, not a diagnosed anxiety population; (d) results are reported as p-values, not effect sizes with CIs; (e) an author became an advisor to a wearables company during the study period. [unverified: the exact analysed N, the anxiety-scale effect size and attrition could not be obtained.]

[unverified] No primary meta-analysis of HRV biofeedback or slow-paced breathing for anxiety meeting this review's source standards was located. No tier assigned.

4.9 Overhyped, and the supplements

InterventionWhat the trials actually show
Cannabinoids / CBDThe largest synthesis — 54 trials, 2,477 participants — found no significant effects on anxiety outcomes, alongside all-cause adverse events OR 1.75 (1.25–2.46), NNTH = 7. (Precision note: the null is reported for cannabinoids generally; the abstract does not isolate a CBD-only anxiety stratum, and the full text is paywalled. Treat "CBD specifically" as unverified — but there is certainly no positive signal.)
KavaGenuinely the best-evidenced botanical: 11 trials, 645 participants; on HAM-A, WMD 5.0 points (1.1–8.8), p = 0.01. But withdrawn in Germany in 2002 over idiosyncratic hepatotoxicity, the review predates modern risk-of-bias tooling, and publication bias in a small manufacturer-adjacent literature is untested. Moderate efficacy, unresolved rare-but-severe hepatic risk.
Silexan (lavender oil)Real trial evidence, N = 539: HAM-A reduction 14.1 (160mg) / 12.8 (80mg) / 11.3 paroxetine / 9.5 placebo, both doses p < 0.01. Two red flags. Two authors are employees of the manufacturer. And assay sensitivity failed — the active comparator paroxetine did not separate from placebo (p = 0.10). A trial where the established drug fails but the sponsor's product succeeds should lower, not raise, your confidence.
AshwagandhaThe clearest example of a broken literature in this review. Four meta-analyses of substantially overlapping trials report incompatible pooled estimates: SMD −6.87, SMD −1.55 (I² = 93.8%, GRADE LOW), HAM-A −5.96 (I² = 98%), HAM-A −2.19. An SMD of −6.87 is nearly seven standard deviations — larger than any psychiatric intervention ever demonstrated, and physically implausible. Threefold spread from the same trials, I² of 93–98%, uniform positivity. Hepatotoxicity is verified: a scoping review of 25 patients found predominantly cholestatic/mixed injury, with one acute liver failure requiring transplantation and three deaths.
Magnesium15 studies, 5 of 7 anxiety studies nominally positive — but narrative only, no meta-analysis, small samples, heterogeneous doses, some with co-ingredients. Serum magnesium shows no association with panic or GAD. Plausible only in baseline deficiency.
Chamomile, weighted blankets, L-theanineNo adequate anxiety-disorder evidence base. L-theanine's verified signal is acute cognitive (choice reaction time after a single 200mg dose), not anxiolytic.

A cross-review note: the stress companion reached the same verdict on ashwagandha from an independent literature, adding that the one review reporting both outcomes found cortisol fell while perceived stress was flatly null (p = 0.40). The biomarker moves; the symptom does not.

4.10 Emerging and contested

Psilocybin for end-of-life anxiety. Two 2016 crossover trials: n = 29 cancer patients (psilocybin vs niacin active placebo) and n = 51 (very low dose 1–3 mg vs high dose 22–30 mg), both with ~60–80% maintaining clinically significant reductions at 6–6.5 months.

The blinding problem is not a footnote; it is the central threat to validity. Niacin produces flushing; 1 mg of psilocybin produces nothing resembling 30 mg. Functional unblinding is essentially certain in both trials. Both mediate the outcome through the mystical-type experience — a self-report measure available only to the unblinded high-dose arm. Combined n across both studies is 80. Promising; not established.

MDMA-assisted therapy. Phase 3 trials reported ~70% of participants no longer meeting PTSD criteria. In 2024 the FDA declined approval, citing blinding failure, absent QT-prolongation and abuse-liability assessments, plus allegations of potential misconduct. No approved MDMA-assisted therapy exists. (Note this is PTSD — no longer an anxiety disorder in DSM-5.)

Ketamine. The strong evidence is for major depressive episodes and suicidality, not anxiety disorders. No adequately powered meta-analysis of ketamine for diagnosed anxiety disorders was located. The only placebo-controlled RCT in a primary anxiety disorder is n = 18, with a broken blind (17/18 patients correctly identified ketamine) and a null co-primary outcome.

4.11 Relapse, non-response, and what "recovery" means

This section is the anxiety counterpart to the stress review's seven-year exhaustion data, and it is less bleak but more equivocal.

  • Response to first-line treatment: 45–65% across anxiety disorders, synthesised from 403 RCTs. Roughly a third to a half of patients do not respond to first-line treatment at all. This number is almost never stated in patient-facing material.
  • Response is not remission. In one GAD trial, 60.3% achieved ≥50% symptom reduction on the best-performing arm but only 46.3% reached the remission threshold — and 29.6% of the placebo arm reached it too.
  • Medication relapse on discontinuation: 36.4% vs 16.4% continued, at up to one year. No data beyond one year.
  • CBT relapse: reported as 0–14% at 3–12 months, but in only 6 of 69 trials. Treat the low figure with suspicion.
  • Durability is the main axis on which psychotherapy beats pharmacotherapy — but that claim rests on k = 3 to k = 10 studies per disorder at ≥12 months, and on zero studies for OCD.

The honest summary for someone starting treatment: the two first-line options are genuinely effective, at roughly half the effect size usually advertised; there is a substantial chance you will be a non-responder; medication gains are largely contingent on continuing it; and whether psychotherapy gains persist past a year is, for most anxiety disorders, unstudied rather than established.

Part V — Synthesis

5.1 The shape of the problem

Four domains, researched independently, produce the same arc so consistently that it is worth treating as the field's central finding rather than as a series of accidents.

The arc: a striking small-sample result → a mechanistic theory built on it → an intervention built on the theory → and then one of three endings.

EndingExamples
A well-powered nullInteroceptive accuracy (55 studies); HPA reactivity (AUCi −0.101, p = 0.28); ENIGMA structural in GAD ("no effect on brain structure"); probiotics in clinical samples; GABA on MRS (p = 0.52); volatility learning (8 experiments, N = 820); ABM registered replication
A confound that accounts for the effectAntidepressants for HRV (unmedicated patients = controls); depression comorbidity for CRP; respiration for HF-HRV; neuroticism for intolerance of uncertainty; comorbidity for days-out-of-role (crude 3× adjusted); familial confounding for attainment (HR 1.56 → 0.96 in discordant twins)
A direct reversalSerotonin up in social anxiety; oestradiol impairing extinction; CRP protective on Mendelian randomisation; Ce and BST responding to the opposite threat types; worry raising rather than suppressing autonomic arousal; rhodiola and placebo (stress companion)

Three structural causes, in order of importance.

1. Measurement. This is the deepest and the most specific to anxiety. The dot-probe has an internal reliability indistinguishable from zero across 36 variants and 9,600 participants. Task-fMRI averages ICC = 0.40 in exactly the regions used as biomarkers. Heartbeat counting correlates r = 0.16 with actual heartbeats and .80–.97 with the number the person says. Cortisol carries 25% between-country variance — more than the entire pooled anxiety effect. You cannot build an individual-differences science on measures that cannot rank individuals, and much of this field tried.

2. Comparator choice. Halving is the rule, not the exception: CBT −1.19 → −0.56, SSRIs −0.91 → −0.44, apps 0.28 → 0.10, ACT significant → null. Over 80% of CBT anxiety trials used waitlists; 17.4% were high quality.

3. The category. rg = 0.91 with depression. GAD structurally clusters with depression, not with panic or the phobias. Intolerance of uncertainty is transdiagnostic and dissolves into neuroticism. The general psychopathology factor performs no better than a raw count of comorbid diagnoses. It is at least as likely that the category is wrong as that the instruments are. You will not find the biology of a thing that is not a kind.

5.2 What the two reviews jointly imply

The stress review and this one were built the same way on different literatures. Three findings survive both, and they are more interesting together than apart.

First, and most substantively: effort is a price, not a fuel gauge — and two independent literatures found it. The stress review reached this from bioenergetics (chronic stress is hypermetabolic), from the death of ego depletion (d = 0.04 and 0.06 across 59 labs), and from Wiehler's lateral-PFC glutamate accumulation. The anxiety literature reached it from Eysenck's attentional control theory, confirmed across 58 studies and 8,292 participants: efficiency falls, effectiveness holds, compensatory effort makes up the difference. Neither field cites the other. That is the strongest available reason to treat the effort-cost framework as a general account rather than a local finding.

Second, the biomarker story fails identically in both. Stress: no cortisol signature of burnout across two systematic syntheses nine years apart; "adrenal fatigue" refuted; flattened diurnal slope predicts health at r = 0.147. Anxiety: cortisol reactivity null, CAR r = .03, no CAR meta-analysis in anxiety disorders even exists. In both fields, the hormone everyone measures is not the variable that predicts the outcome anyone cares about — and in both, that has not stopped a commercial market from selling the measurement.

Third, the intervention hierarchy is nearly the same, and it is unglamorous. Structured behavioural protocols and, where indicated, medication. Then exercise, with honest comparators. Then a long tail of things that vanish against active controls: mindfulness apps (0.09 for stress, 0.10 for anxiety, both non-significant), adaptogens, supplements in non-deficient people. Ashwagandha fails identically in both reviews from independent literatures.

Where they genuinely diverge, and it matters clinically: the cognitive signature. Anxiety costs more to run; burnout runs worse. Someone who cannot face their work may be paying a raised price with intact machinery, or working with degraded machinery — and the two call for opposite responses. Neither is diagnosable from self-report, because in both conditions subjective and objective impairment come apart.

5.3 Five open questions worth a research programme

1. Can anything in this field be measured reliably enough for individual-differences work? This is prior to every other question. Trial-level and dynamic scoring of the dot-probe reportedly shows good reliability where difference scores show none; that lead has not been properly pursued. The discriminating study is not another bias–symptom correlation but a psychometric one: establish which anxiety-relevant measures clear ICC ≈ 0.7, then rebuild the individual-differences literature only on those. Everything downstream depends on it.

2. Is the fear/distress carve better than the DSM's? Structural evidence puts GAD with depression and panic with the phobias. The manuals disagree. This is testable and rarely tested: do treatment response, course, and genetic architecture track the DSM category or the fear/distress split? If the latter, the practical consequence is large — it would mean GAD trials and panic trials should not be pooled, which they routinely are.

3. Why does normalising the physiology not normalise the experience? Three independent demonstrations exist: CRF1 antagonism enhanced startle inhibition while failing clinically; 90% receptor occupancy produced no anxiolysis; hypoventilation therapy normalised PCO₂ with no differential symptom benefit. Either the physiological targets are epiphenomenal, or the subjective state has a separate cortical substrate — LeDoux's claim, largely untested as an empirical prediction rather than a philosophical position. The discriminating experiment is a within-subject design measuring circuit engagement and subjective state independently across a manipulation that moves one without the other.

4. Does prevalence inflation actually happen? Foulkes & Andrews' hypothesis is three years old, theoretically well-motivated — the proposed mechanism is avoidance, the best-supported maintaining factor in the field — and has never been empirically tested. The design is not exotic: randomise exposure to framing that labels mild distress as a disorder versus framing that normalises it, follow avoidance behaviour and symptoms prospectively. Given that the answer bears on how every school and employer in the developed world runs its mental-health messaging, its absence from the literature is remarkable.

5. Do CBT gains hold past a year? For most anxiety disorders this is unstudied, not established: k = 3 for social anxiety, k = 1 for specific phobia, k = 0 for OCD at ≥12 months, and non-significant for panic. Relapse was reported in 6 of 69 trials. This is the cheapest high-value study in the whole field — it requires no new treatment, only following existing cohorts and reporting what happens.

5.4 Practical translation

Nothing here is medical advice. What follows is what the evidence supports, stated at the level of usefulness.

The two things that work

Individual CBT with exposure, and SSRIs/SNRIs. They were the only two intervention classes to beat an appropriate placebo, at roughly −0.5 SMD each — real, moderate, and about half what waitlist-controlled trials advertise. Guided internet CBT is equivalent to face-to-face (g = 0.02), which makes access less of a barrier than it was; unguided is substantially worse, mostly via completion (48% vs 79%).

Expect a 45–65% response rate. A third to a half of people do not respond to first-line treatment, and knowing that in advance is better than discovering it and concluding you are the problem.

What to be sceptical of

  • Anything sold on a biomarker. Salivary cortisol panels, HRV-based "nervous system" assessments, and gut-microbiome anxiety testing have no validated basis. HRV differences in anxiety are substantially a medication effect; cortisol reactivity is null; microbiome evidence in anxiety disorders totals three studies and 84 patients.
  • Apps as treatment. Against an active or placebo app, the anxiety effect is g = 0.10, non-significant. Against nothing, 0.28. The difference is what expectancy and attention buy.
  • Supplements. Ashwagandha's literature contains no negative trials and pooled estimates spanning threefold, with verified hepatotoxicity including deaths. CBD shows no anxiety signal in 54 trials with adverse events at NNTH 7. Magnesium has no study using a validated stress measure. Kava is the one botanical with a real signal, and it carries a rare hepatic risk that got it withdrawn in Germany.
  • Propranolol for performance anxiety. Eight trials total across all anxiety disorders; the largest social-anxiety study had 16 participants. It may work. Nobody knows.
  • "Push through it, the arousal helps." Yerkes–Dodson is forty mice and two per data point in the only condition that produced the curve. Facilitative anxiety, on meta-analysis, is self-confidence wearing a different name (cognitive anxiety r = −0.10; self-confidence r = +0.24).

Two things worth internalising

Reassurance-seeking does not work, and this is measurable. Across 14 RCTs and 3,828 patients, diagnostic testing to reassure produced null effects on illness worry, anxiety and symptom persistence. A negative scan does not settle the question, because the question was never really about the scan. This is one of the most practically useful findings in the review and one of the least known.

Avoidance is the mechanism worth watching. It is the common thread in every cognitive-behavioural model, the target of the one treatment class that beats placebo, and — if the prevalence inflation hypothesis is right — the mechanism by which labelling mild anxiety as disorder could make it worse. The evidence for avoidance as the driver of impairment is thinner than the theory implies (no well-powered study pits it against arousal). But it is the thing that both the working treatments and the best causal hypotheses converge on.

A note on the category

If you are reading this because the label matters to you personally: the genetic correlation between anxiety and depression is 0.91, GAD sits structurally closer to depression than to panic, and 63% of people with a current anxiety disorder currently also have depression. The boundaries are administratively real and biologically blurry. That is not a reason to distrust a diagnosis — it can still be the right entry point to effective treatment — but it is a reason not to over-invest in which side of a line you fall on, or to conclude that a treatment developed for the neighbouring category cannot help.

Appendix A — Key numbers

FindingValueSource
MECHANISMS
Ce vs BST response to threatStatistically indistinguishable, Bayesian equivalence, N = 295Didier et al. 2026, SCAN 21:nsag040
Task-fMRI test–retest reliabilitymean ICC = 0.397 (90 experiments, N = 1,008)Elliott et al. 2020, Psychol Sci 31:792–806
Largest replicating brain-wide association|r| = 0.16; 80% power needs n = 9,500Marek et al. 2022, Nature 603:654–660
ENIGMA GAD (1,020 vs 2,999)"No effect of GAD on brain structure"Harrewijn et al. 2021, Transl Psychiatry 11:502
ENIGMA social anxiety, putamend = −0.077 left / −0.104 rightGroenewold et al. 2023, Mol Psychiatry 28:1079
ENIGMA panic, case-controld = −0.07 to −0.13; but early-onset lateral ventricles d = 0.31–0.38Han et al. 2026, Mol Psychiatry
Cardiac interoceptive accuracy × anxietyNo association (55 studies, pre-registered)Adams et al. 2022, Neurosci Biobehav Rev 140:104754
Heartbeat counting: actual vs reportedr = 0.16; scores correlate .80–.97 with raw reported countZamariola et al. 2018, Biol Psychol 137:12–17
HF-HRV, all anxiety disordersg = −0.29 (−0.41, −0.17)Chalmers et al. 2014, Front Psychiatry 5:80
Time-domain HRV before outlier deletiong = −0.70 (−1.45, −0.05), p = 0.07, I² = 98.2%Chalmers et al. 2014, Table 2
Unmedicated anxious patients vs controls, HRVNo differenceLicht et al. 2009, Psychosom Med 71:508
Cortisol stress reactivity (AUCi), anxiety−0.101 (−0.285, 0.082), p = 0.28Zorn et al. 2017, Psychoneuroendocrinology 77:25
Cortisol awakening responser = .03, p = .13 (k = 59)Boggero et al. 2017, Biol Psychol 129:207
TSST variance attributable to country~25%; N America d = 0.45 vs Europe d = 0.73Miller & Kirschbaum 2019
MRS GABA in anxietyg = −0.26, p = 0.52, k = 3, n = 42Maddock & Smucny 2025, Mol Psychiatry 30:6020
Drug-minus-placebo, anxiety disordersg = 0.26–0.39 (placebo pre-post g = 0.70–1.10)Sugarman et al. 2017
Twin heritability, panic / GAD0.48 (0.41–0.54) / 0.316 (0.24–0.39)Hettema et al. 2001, Am J Psychiatry 158:1568
SNP heritability, largest GWAS10.1% — ~¼ of twin h²; 122,341 casesStrom et al. 2026, Nat Genet 58:275–288
Genetic correlation with depressionrg = 0.91 (range 0.78–0.91)Strom et al. 2026
Polygenic score variance explained2.27% EUR; 0.54% AFR (p = 0.051, ns)Strom et al. 2026
5-HTTLPR × stress interactionOR 1.01 (0.94–1.10), N = 14,250Risch et al. 2009, JAMA 301:2462
CRP × anxiety after adjusting for depressionOR 0.98 (0.97–0.99) — reversesYe et al. 2021, eClinicalMedicine 38:100992
CRP → anxiety, Mendelian randomisationOR 0.87 (0.80–0.95) — protectiveYe et al. 2021
Probiotics: animals vs humans, same paperg = −0.47 (p = .004) vs −0.12 (p = .151), nullReis et al. 2018, PLoS One 13:e0199041
Microbiome studies in anxiety disorders3 studies, 84 patients (and 96.6% positivity across all disorders)Nikolova et al. 2021, JAMA Psychiatry 78:1343
Sex ratio, any anxiety disorderOR 1.7 (1.6–1.8) lifetime; crude prevalence ratio 1.51Seedat et al. 2009; McLean et al. 2011
Oestradiol × fear extinction, best-powered RCTImpaired extinction recall, N = 116Kaczmarczyk et al. 2024, Transl Psychiatry 14:437
GAD-7 measurement invariance across sexFull invariance, N = 165,872 (largest ΔCFI −0.006)Saunders et al. 2023, BMC Psychiatry 23:301
Elevated plus maze measures reliably detecting anxiolytics2 of 17 (814 studies, pre-registered)Rosso et al. 2022, Neurosci Biobehav Rev 143:104928
Single-lab mouse behavioural discovery, P(non-replicable)59.6%Jaljuli et al. 2023, PLoS Biol 21:e3002082
COGNITION
Threat attentional bias, anxious vs non-anxiousd = 0.45; 172 studiesBar-Haim et al. 2007, Psychol Bull 133:1–24
Dot-probe internal reliability, 36 variantsNone greater than zero; N = 9,600Xu et al. 2025, Clin Psychol Sci 13:261–277
ABM, clinical general anxietyg = 0.29 (−0.14 to 0.74), nsCristea et al. 2015, BJPsych 206:7–16
ABM registered replicationBayesian null on 4 of 5 measures, incl. the bias; N = 104Pond et al. 2025, Collabra 11:147691
Interpretation bias modification vs shamSMD −0.30 (−0.50 to −0.10); prediction intervals include 0Fodor et al. 2020, Lancet Psychiatry
Intolerance of uncertainty, pooledr = 0.51 (0.50–0.52); 181 studies, N = 52,402 — transdiagnosticMcEvoy et al. 2019
IU after controlling neuroticism facetsNo unique association with any disorderNaragon-Gainey & Watson 2018, Assessment 25:143
Perseverative cognition → SBP / DBP / HR / cortisolg = .45 / .51 / .28 / .36 — all activatingOttaviani et al. 2016, Psychol Bull 142:231
Anxiety × attentional controlg = −0.58; efficiency not effectiveness; spares updatingShi et al. 2019, Clin Psychol Rev 72:101754
Anxiety × working memoryg = −0.334; 177 samples, N = 22,061Moran 2016, Psychol Bull 142:831
Volatility learning, largest testNo systematic effect; 8 experiments, N = 820 (preprint)Satti et al. 2025
Safety-behaviour elimination vs judicious useNo group differences, N = 60 and N = 56Blakey et al. 2019; Meckes et al. 2025
Yerkes & Dodson 190840 mice; 2 per data point in the critical condition; easy condition monotonic; zero statistics; "arousal" appears 0 timesOriginal text
Competitive cognitive anxiety × performancer = −0.10 (self-confidence r = +0.24); k = 48Woodman & Hardy 2003
Math anxiety × performancer = −.30, N = 906,311 (range −.168 to −.34 across meta-analyses)Caviola et al. 2022
Working-memory mediation of math anxietyIndirect effects all ns, B ≤ .05Caviola et al. 2022
Math anxiety ⇄ achievement, longitudinalMA→MP b = −.11; MP→MA b = −.12; N = 39,935Pelegrina et al. 2025
Stereotype threat, girls and mathsg = −0.22 → trim-fill −0.07, p = .27; operational settings d = −.01Flore & Wicherts 2015; Shewach et al. 2019
Anxiety sensitivity → panic (5 weeks)Upper ASI decile 20% vs 6%; N = 1,401Schmidt et al. 1997
Anxiety sensitivity reduction, long-termd = 0.29, nsFitzgerald et al. 2021
FUNCTION & EPIDEMIOLOGY
Adjusted additional days out of rolePanic 14.3, PTSD 15.2, GAD 7.7 (crude figures ~3× larger)Alonso et al. 2011, Mol Psychiatry 16:1234
Reassurance from negative diagnostic testingIllness worry OR 0.87 (0.55–1.39) — nullRolfe & Burton 2013, JAMA Intern Med 173:407
Anxiety disorders → motor-vehicle crashesNo consistent association; objective-measure study OR 0.64/0.69 nsRapoport et al. 2023
Benzodiazepines → crash riskOR 1.59 (1.10–2.31); with alcohol 7.69 (4.33–13.65)Dassanayake et al. 2011
Educational attainment, GAD/social phobia/PTSDNon-significant at 12 of 12 milestonesBreslau et al. 2008
Internalizing → unemployment, twin-discordantHR 1.56 → 0.96 (0.57–1.62), p = .88Alaie et al. 2023
Global anxiety prevalence, adjusted7.3% (4.8–10.9)pointBaxter et al. 2013, Psychol Med 43:897
Global anxiety, GBD 20193.78%point; 301.4M casesGBD 2019, Lancet Psychiatry 9:137
Cross-national anxiety, 21 countries9.8% (3.0–19.0%) — 12-monthAlonso et al. 2018
Trend 1990 → 20103.8% → 4.0% age-standardised — flatBaxter et al. 2014, Depress Anxiety 31:506
UK youth anxiety diagnosis, 2003–2018IRR 3.51 (3.18–3.89)Cybulski et al. 2021
UK anxiety diagnosis incidence, April 2020−47.8% in one monthCarr et al. 2021, Lancet Public Health 6:e124
Youth anxiety, England 2004 → 2017, same instrument3.5% → 5.4%, RR 1.63 (1.37–1.93)Taxiarchi et al. 2026, BJPsych 229:36
COVID modelled increase+25.6% (23.2–28.0), from 27 anxiety studiesSantomauro et al. 2021, Lancet 398:1700
COVID observed change, same-participant cohortsSMD 0.05 (−0.04 to 0.13), nsSun et al. 2023, BMJ 380:e074224
Age at onset, anxiety block vs GADBlock peak 5.5y; GAD median 30–35y; panic 25–27ySolmi et al. 2022, Mol Psychiatry 27:281
Current anxiety disorder with comorbid depression63% current, 81% lifetimeLamers et al. 2011
12-year recovery / recurrence, social phobia0.37 / 0.39Bruce et al. 2005, Am J Psychiatry 162:1179
Chronicity at 2 years, pure anxiety41.9% chronic, median 16 monthsPenninx et al. 2011
Recurrence, index-only vs any new disorder23.8% → 54.8% over 4 yearsScholten et al. 2016
Anxiety → depression / depression → anxietyOR 2.77 vs 2.73 — symmetricalJacobson & Newman 2017, Psychol Bull 143:1155
Incident coronary heart diseaseHR 1.26 (1.15–1.38)Roest et al. 2010, JACC 56:38
All-cause mortality1.09 → 1.03 (0.95–1.13) bias-adjusted; 0.99 communityMiloyan et al. 2016
Any treatment / possibly adequate treatment27.6% / 9.8%; perceived need only 41.3%Alonso et al. 2018
INTERVENTIONS
Individual CBT, social anxiety−1.19 waitlist → −0.56 psychological placeboMayo-Wilson et al. 2014, Lancet Psychiatry 1:368
SSRIs/SNRIs, social anxiety−0.91 waitlist → −0.44 pill placeboMayo-Wilson et al. 2014
CBT vs placebo, anxiety-related disordersg = 0.56; response OR 2.97Carpenter et al. 2018
CBT trials rated high quality17.4%; >80% used waitlist controlsCuijpers et al. 2016, World Psychiatry 15:245
CBT at ≥12 monthsGAD g = 0.22; panic ns; specific phobia k = 1; OCD k = 0van Dis et al. 2020, JAMA Psychiatry 77:265
Antidepressants, panic disorderRR 0.72 (0.66–0.79), NNTB 7 (6–9)Bighelli et al. 2018, Cochrane
GAD pharmacotherapy, best agentsQuetiapine −3.60; duloxetine −3.13; pregabalin −2.79 HAM-ASlee et al. 2019, Lancet 393:768
Antidepressant discontinuation symptoms31% vs 17% placebo; ~15% net; severe 2.8% vs 0.6%Henssler et al. 2024, Lancet Psychiatry 11:526
Relapse on discontinuation, 1 year36.4% vs 16.4% continued; OR 3.11Batelaan et al. 2017, BMJ 358:j3927
Response rate, first-line anxiety treatment45–65% (403 RCTs)Bandelow et al. 2014
Guided iCBT vs face-to-face CBTg = 0.02 (−0.09 to 0.14) — equivalentHedman-Lagerlöf et al. 2023, World Psychiatry 22:305
Completion: face-to-face / guided / self-guided86.7% / 79.1% / 48.2%Bakanaitė et al. 2025
Apps for anxiety vs active/placebo appg = 0.10, ns (vs passive 0.28)Linardon et al. 2024, Clin Psychol Rev 107:102370
Exercise, diagnosed anxiety disordersSMD −0.582, 6 RCTs, n = 262, CI misprintedStubbs et al. 2017
Mindfulness vs active controls, anxietyES 0.38 (0.12–0.64); no better than any active treatmentGoyal et al. 2014, JAMA Intern Med 174:357
MBSR vs escitalopramNoninferior, difference −0.07 (−0.38 to 0.23); AEs 15.4% vs 78.6%Hoge et al. 2023, JAMA Psychiatry 80:13
MBSR videoconference replicationNoninferiority NOT supported, p = 0.17Hoge et al. 2025, J Affect Disord 384:163
Applied relaxation, GAD, low-RoB only−0.47 (−1.18 to 0.23) — loses significancePapola et al. 2024, JAMA Psychiatry 81:250
D-cycloserine augmentation at follow-upd = −0.19, CI crosses zeroMataix-Cols et al. 2017, JAMA Psychiatry 74:501
Propranolol: total RCT evidence, all anxiety disorders8 studies; largest social phobia trial n = 16Steenen et al. 2016, J Psychopharmacol 30:128
Cannabinoids for anxietyNo significant effect; AEs OR 1.75, NNTH 7; 54 trialsWilson et al. 2026, Lancet Psychiatry 13:304
KavaHAM-A WMD 5.0 (1.1–8.8); withdrawn in Germany 2002Pittler & Ernst 2003, Cochrane
Silexan trial: paroxetine active comparatorFailed to separate from placebo (p = 0.10)Kasper et al. 2014
Ashwagandha pooled estimates, same trialsSMD −1.55 to −6.87, I² 93–98%Four meta-analyses, 2022–2026
Ashwagandha liver injury25 patients; 1 transplant, 3 deathsMcIntyre et al. 2026
CBT + taper vs taper alone, benzodiazepine cessationNNT 2.8 (1.9–5.3) at 6–12 monthsTakeshima et al. 2021
Psilocybin, end-of-life anxietyn = 29 and n = 51; blinding compromised in bothRoss 2016; Griffiths 2016

Appendix B — Master reference list

Mechanisms

  1. Didier PR, Grogans SE, … Shackman AJ (2026). Fear, anxiety, and the extended amygdala — absence of evidence for strict functional segregation. Soc Cogn Affect Neurosci 21:nsag040. https://pubmed.ncbi.nlm.nih.gov/42234867/
  2. Davis M, Walker DL, Miles L, Grillon C (2010). Phasic vs sustained fear in rats and humans. Neuropsychopharmacology 35:105–135. https://pubmed.ncbi.nlm.nih.gov/19693004/
  3. Shackman AJ, Fox AS (2016). Contributions of the central extended amygdala to fear and anxiety. J Neurosci 36:8050–8063. https://pubmed.ncbi.nlm.nih.gov/27488625/
  4. LeDoux JE, Pine DS (2016). Using neuroscience to help understand fear and anxiety: a two-system framework. Am J Psychiatry 173:1083–1093. https://pubmed.ncbi.nlm.nih.gov/27609244/
  5. Fanselow MS, Pennington ZT (2018). A return to the psychiatric dark ages with a two-system framework for fear. Behav Res Ther 100:24–29. https://pubmed.ncbi.nlm.nih.gov/29128585/
  6. Mobbs D, Adolphs R, Fanselow MS, et al. (2019). Viewpoints: approaches to defining and investigating fear. Nat Neurosci 22:1205–1216. https://pubmed.ncbi.nlm.nih.gov/31332374/
  7. Etkin A, Wager TD (2007). Functional neuroimaging of anxiety. Am J Psychiatry 164:1476–1488. https://pubmed.ncbi.nlm.nih.gov/17898336/
  8. Chavanne AV, Robinson OJ (2021). The overlapping neurobiology of induced and pathological anxiety. Am J Psychiatry 178:156–164. https://pubmed.ncbi.nlm.nih.gov/33054384/
  9. Elliott ML, et al. (2020). What is the test–retest reliability of common task-fMRI measures? Psychol Sci 31:792–806. https://pubmed.ncbi.nlm.nih.gov/32489141/
  10. Marek S, et al. (2022). Reproducible brain-wide association studies require thousands of individuals. Nature 603:654–660. https://pubmed.ncbi.nlm.nih.gov/35296861/
  11. Harrewijn A, et al. (2021). Cortical and subcortical brain structure in generalized anxiety disorder: ENIGMA-Anxiety. Transl Psychiatry 11:502. https://pubmed.ncbi.nlm.nih.gov/34599145/
  12. Groenewold NA, et al. (2023). Volume of subcortical brain regions in social anxiety disorder. Mol Psychiatry 28:1079–1089. https://pubmed.ncbi.nlm.nih.gov/36653677/
  13. Han S, et al. (2026). Brain structure in panic disorder: ENIGMA-Anxiety. Mol Psychiatry. https://pubmed.ncbi.nlm.nih.gov/41530384/
  14. Adams KL, et al. (2022). Interoception and social anxiety: a systematic review and meta-analysis. Neurosci Biobehav Rev 140:104754. https://pubmed.ncbi.nlm.nih.gov/35798125/
  15. Zamariola G, et al. (2018). Interoceptive accuracy scores from the heartbeat counting task are problematic. Biol Psychol 137:12–17. https://pubmed.ncbi.nlm.nih.gov/29944964/
  16. Chalmers JA, et al. (2014). Anxiety disorders are associated with reduced heart rate variability. Front Psychiatry 5:80. https://pubmed.ncbi.nlm.nih.gov/25071612/
  17. Licht CM, et al. (2009). Association between anxiety disorders and heart rate variability. Psychosom Med 71:508–518. https://pubmed.ncbi.nlm.nih.gov/19414616/
  18. Menuet C, et al. (2025). Expert recommendations on respiratory sinus arrhythmia. Nat Rev Cardiol. https://pubmed.ncbi.nlm.nih.gov/40328963/
  19. Zorn JV, et al. (2017). Cortisol stress reactivity across psychiatric disorders. Psychoneuroendocrinology 77:25–36. https://pubmed.ncbi.nlm.nih.gov/28012291/
  20. Miller R, Kirschbaum C (2019). Cultures under stress: a cross-national meta-analysis of cortisol responses. Neurosci Biobehav Rev. https://pubmed.ncbi.nlm.nih.gov/30611610/
  21. Maddock RJ, Smucny J (2025). Magnetic resonance spectroscopy in anxiety disorders: a meta-analysis. Mol Psychiatry 30:6020–6032. https://pubmed.ncbi.nlm.nih.gov/40913113/
  22. Moncrieff J, et al. (2022). The serotonin theory of depression: a systematic umbrella review. Mol Psychiatry 28:3243–3256. https://pubmed.ncbi.nlm.nih.gov/35854107/
  23. Jauhar S, et al. (2023). A leaky umbrella has little value: evidence clearly favours the serotonin hypothesis. Mol Psychiatry 28:3149–3152. https://pubmed.ncbi.nlm.nih.gov/37322065/
  24. Frick A, et al. (2015). Serotonin synthesis and reuptake in social anxiety disorder. JAMA Psychiatry 72:794–802. https://pubmed.ncbi.nlm.nih.gov/26083190/
  25. Strom NI, et al. (2026). Genome-wide association study of anxiety. Nat Genet 58:275–288. https://pmc.ncbi.nlm.nih.gov/articles/PMC12900644/
  26. Hettema JM, Neale MC, Kendler KS (2001). A review and meta-analysis of the genetic epidemiology of anxiety disorders. Am J Psychiatry 158:1568–1578. https://pubmed.ncbi.nlm.nih.gov/11578982/
  27. Border R, et al. (2019). No support for historical candidate gene or candidate gene-by-interaction hypotheses for major depression. Am J Psychiatry 176:376–387. https://pubmed.ncbi.nlm.nih.gov/30845820/
  28. Ye Z, et al. (2021). Role of inflammation in depression and anxiety: UK Biobank and Mendelian randomisation. eClinicalMedicine 38:100992. https://pubmed.ncbi.nlm.nih.gov/34505025/
  29. Reis DJ, Ilardi SS, Punt SEW (2018). The anxiolytic effect of probiotics: a systematic review and meta-analysis. PLoS One 13:e0199041. https://pubmed.ncbi.nlm.nih.gov/29924822/
  30. Nikolova VL, et al. (2021). Perturbations in gut microbiota composition in psychiatric disorders. JAMA Psychiatry 78:1343–1354. https://pubmed.ncbi.nlm.nih.gov/34524405/
  31. Seedat S, et al. (2009). Cross-national associations between gender and mental disorders. Arch Gen Psychiatry 66:785–795. https://pubmed.ncbi.nlm.nih.gov/19581570/
  32. Kaczmarczyk M, et al. (2024). Estradiol and fear extinction: a randomised controlled trial. Transl Psychiatry 14:437. https://pubmed.ncbi.nlm.nih.gov/39448569/
  33. Rosso M, et al. (2022). Reliability of common mouse behavioural tests of anxiety. Neurosci Biobehav Rev 143:104928. https://pubmed.ncbi.nlm.nih.gov/36341943/
  34. Bach DR (2022). Cross-species anxiety tests in psychiatry: pitfalls and promises. Mol Psychiatry 27:154–163. https://pubmed.ncbi.nlm.nih.gov/34561614/

Cognition and function

  1. Bar-Haim Y, et al. (2007). Threat-related attentional bias in anxious and nonanxious individuals: a meta-analytic study. Psychol Bull 133:1–24. https://pubmed.ncbi.nlm.nih.gov/17201568/
  2. Xu I, Passell E, Strong RW, et al. (2025). No evidence of reliability across 36 variations of the emotional dot-probe task. Clin Psychol Sci 13:261–277. https://pubmed.ncbi.nlm.nih.gov/40151297/
  3. Rodebaugh TL, et al. (2016). Unreliability as a threat to understanding psychopathology. J Abnorm Psychol 125:840–851. https://pubmed.ncbi.nlm.nih.gov/27322741/
  4. Cristea IA, Kok RN, Cuijpers P (2015). Efficacy of cognitive bias modification interventions in anxiety and depression. Br J Psychiatry 206:7–16. https://pubmed.ncbi.nlm.nih.gov/25561486/
  5. Fodor LA, et al. (2020). Efficacy of cognitive bias modification interventions in anxiety and depressive disorders: network meta-analysis. Lancet Psychiatry 7:506–514. https://pubmed.ncbi.nlm.nih.gov/32445689/
  6. Pond RM, Meeten F, Clarke PJF, Notebaert L, Scott R (2025). A registered replication of attention bias modification in high worriers. Collabra: Psychology 11:147691. https://doi.org/10.1525/collabra.147691
  7. Newman MG, Llera SJ (2011). A novel theory of experiential avoidance in GAD. Clin Psychol Rev 31:371–382. https://pmc.ncbi.nlm.nih.gov/articles/PMC3073849/
  8. Ottaviani C, et al. (2016). Physiological concomitants of perseverative cognition: a meta-analysis. Psychol Bull 142:231–259. https://pubmed.ncbi.nlm.nih.gov/26689087/
  9. McEvoy PM, et al. (2019). Intolerance of uncertainty and negative metacognitive beliefs as transdiagnostic mediators. Clin Psychol Rev 73:101778. https://pubmed.ncbi.nlm.nih.gov/31678816/
  10. Naragon-Gainey K, Watson D (2018). What lies beyond neuroticism? Assessment 25:143–158. https://pubmed.ncbi.nlm.nih.gov/27411679/
  11. Eysenck MW, Derakshan N, Santos R, Calvo MG (2007). Anxiety and cognitive performance: attentional control theory. Emotion 7:336–353. https://pubmed.ncbi.nlm.nih.gov/17516812/
  12. Shi R, Sharpe L, Abbott M (2019). A meta-analysis of the relationship between anxiety and attentional control. Clin Psychol Rev 72:101754. https://pubmed.ncbi.nlm.nih.gov/31306935/
  13. Moran TP (2016). Anxiety and working memory capacity: a meta-analysis and narrative review. Psychol Bull 142:831–864. https://pubmed.ncbi.nlm.nih.gov/26963369/
  14. Browning M, Behrens TE, Jocham G, O'Reilly JX, Bishop SJ (2015). Anxious individuals have difficulty learning the causal statistics of aversive environments. Nat Neurosci 18:590–596. https://pubmed.ncbi.nlm.nih.gov/25730669/
  15. Gagne C, Zika O, Dayan P, Bishop SJ (2020). Impaired adaptation of learning to contingency volatility in internalizing psychopathology. eLife 9:e61387. https://pubmed.ncbi.nlm.nih.gov/33350387/
  16. Duits P, et al. (2015). Updated meta-analysis of classical fear conditioning in the anxiety disorders. Depress Anxiety 32:239–253. https://pubmed.ncbi.nlm.nih.gov/25703487/
  17. Yerkes RM, Dodson JD (1908). The relation of strength of stimulus to rapidity of habit-formation. J Comp Neurol Psychol 18:459–482. https://psychclassics.yorku.ca/Yerkes/Law/
  18. Woodman T, Hardy L (2003). The relative impact of cognitive anxiety and self-confidence upon sport performance. J Sports Sci 21:443–457. https://pubmed.ncbi.nlm.nih.gov/12846532/
  19. Caviola S, et al. (2022). Math performance and academic anxiety forms: a meta-analytic review. Educ Psychol Rev 34:363–399. https://link.springer.com/article/10.1007/s10648-021-09618-5
  20. Flore PC, Wicherts JM (2015). Does stereotype threat influence performance of girls in stereotyped domains? J Sch Psychol 53:25–44. https://pubmed.ncbi.nlm.nih.gov/25636259/
  21. Shewach OR, Sackett PR, Quint S (2019). Stereotype threat effects in settings with features likely versus unlikely in operational test settings. J Appl Psychol 104:1514–1534. https://pubmed.ncbi.nlm.nih.gov/31094540/
  22. Schmidt NB, Lerew DR, Jackson RJ (1997). The role of anxiety sensitivity in the pathogenesis of panic. J Abnorm Psychol 106:355–364. https://pubmed.ncbi.nlm.nih.gov/9241937/
  23. Alonso J, et al. (2011). Days out of role due to common physical and mental conditions. Mol Psychiatry 16:1234–1246. https://pmc.ncbi.nlm.nih.gov/articles/PMC3223313/
  24. Rolfe A, Burton C (2013). Reassurance after diagnostic testing with a low pretest probability of serious disease. JAMA Intern Med 173:407–416. https://pubmed.ncbi.nlm.nih.gov/23440131/
  25. Rapoport MJ, et al. (2023). Psychiatric disorders and motor vehicle collision risk. Can J Psychiatry 68:221–240. https://pmc.ncbi.nlm.nih.gov/articles/PMC10037743/
  26. Alaie I, Svedberg P, Ropponen A, Narusyte J (2023). Adolescent internalizing problems and adult labour-market marginalisation: a co-twin control study. JAMA Netw Open 6:e2317905.

Taxonomy, epidemiology and course

  1. Reed GM, First MB, Kogan CS, et al. (2019). Innovations and changes in the ICD-11 classification of mental disorders. World Psychiatry 18:3–19. https://pubmed.ncbi.nlm.nih.gov/30600616/
  2. Stein DJ, Craske MA, Friedman MJ, Phillips KA (2014). Anxiety disorders, OCD and related disorders, and trauma- and stressor-related disorders in DSM-5. Am J Psychiatry 171:611–613. https://pubmed.ncbi.nlm.nih.gov/24880507/
  3. Abramowitz JS, Jacoby RJ (2015). Obsessive-compulsive and related disorders: a critical review of the new diagnostic class. Annu Rev Clin Psychol 11:165–186. https://pubmed.ncbi.nlm.nih.gov/25581239/
  4. Lee S, Tsang A, Ruscio AM, et al. (2009). Implications of modifying the duration requirement of GAD. Psychol Med 39:1163–1176. https://pubmed.ncbi.nlm.nih.gov/19091158/
  5. Ruscio AM, et al. (2017). Cross-sectional comparison of the epidemiology of DSM-5 generalized anxiety disorder across the globe. JAMA Psychiatry 74:465–475. https://pubmed.ncbi.nlm.nih.gov/28297020/
  6. Baxter AJ, Scott KM, Vos T, Whiteford HA (2013). Global prevalence of anxiety disorders: a systematic review and meta-regression. Psychol Med 43:897–910.
  7. Baxter AJ, Scott KM, Ferrari AJ, et al. (2014). Challenging the myth of an "epidemic" of common mental disorders. Depress Anxiety 31:506–516. https://pubmed.ncbi.nlm.nih.gov/24448889/
  8. Santomauro DF, et al. (2021). Global prevalence and burden of depressive and anxiety disorders in 204 countries in 2020. Lancet 398:1700–1712. https://pubmed.ncbi.nlm.nih.gov/34634250/
  9. Sun Y, et al. (2023). Comparison of mental health symptoms before and during the COVID-19 pandemic. BMJ 380:e074224. https://pubmed.ncbi.nlm.nih.gov/36889797/
  10. Taxiarchi VP, et al. (2026). Trends in child and adolescent emotional disorders in England. Br J Psychiatry 229:36–42. https://pubmed.ncbi.nlm.nih.gov/40546093/
  11. Cybulski L, et al. (2021). Temporal trends in annual incidence rates for psychiatric disorders in UK primary care. BMC Psychiatry 21:229. https://pubmed.ncbi.nlm.nih.gov/33941129/
  12. Carr MJ, et al. (2021). Effects of the COVID-19 pandemic on primary care-recorded mental illness. Lancet Public Health 6:e124–e135. https://pubmed.ncbi.nlm.nih.gov/33444560/
  13. Foulkes L, Andrews JL (2023). Are mental health awareness efforts contributing to the rise in reported mental health problems? New Ideas Psychol 69:101010. https://doi.org/10.1016/j.newideapsych.2023.101010
  14. Solmi M, et al. (2022). Age at onset of mental disorders worldwide. Mol Psychiatry 27:281–295. https://pubmed.ncbi.nlm.nih.gov/34079068/
  15. Watson D (2005). Rethinking the mood and anxiety disorders: a quantitative hierarchical model. J Abnorm Psychol 114:522–536. https://pubmed.ncbi.nlm.nih.gov/16351375/
  16. Kotov R, et al. (2017). The Hierarchical Taxonomy of Psychopathology (HiTOP). J Abnorm Psychol 126:454–477. https://pubmed.ncbi.nlm.nih.gov/28333488/
  17. Forbes MK, et al. (2021). On unreplicable inferences in psychopathology research. J Abnorm Psychol 130:297–317. https://pubmed.ncbi.nlm.nih.gov/33539117/
  18. Bruce SE, et al. (2005). Influence of psychiatric comorbidity on recovery and recurrence in GAD, social phobia and panic disorder. Am J Psychiatry 162:1179–1187. https://pmc.ncbi.nlm.nih.gov/articles/PMC3272761/
  19. Penninx BW, et al. (2011). Two-year course of depressive and anxiety disorders: NESDA. J Affect Disord 133:76–85. https://pubmed.ncbi.nlm.nih.gov/21496929/
  20. Scholten WD, et al. (2016). Diagnostic instability of recurrence and the impact on recurrence rates. J Affect Disord 195:185–190. https://pubmed.ncbi.nlm.nih.gov/26896812/
  21. Jacobson NC, Newman MG (2017). Anxiety and depression as bidirectional risk factors for one another. Psychol Bull 143:1155–1200. https://pubmed.ncbi.nlm.nih.gov/28805400/
  22. Roest AM, et al. (2010). Anxiety and risk of incident coronary heart disease. J Am Coll Cardiol 56:38–46. https://pubmed.ncbi.nlm.nih.gov/20620715/
  23. Miloyan B, et al. (2016). Anxiety disorders and all-cause mortality: a meta-analysis. Soc Psychiatry Psychiatr Epidemiol 51:1467–1475. https://pubmed.ncbi.nlm.nih.gov/27628244/
  24. Alonso J, et al. (2018). Treatment gap for anxiety disorders is global. Depress Anxiety 35:195–208. https://pmc.ncbi.nlm.nih.gov/articles/PMC6008788/

Interventions

  1. Cuijpers P, et al. (2016). How effective are cognitive behavior therapies for major depression and anxiety disorders? World Psychiatry 15:245–258. https://pubmed.ncbi.nlm.nih.gov/27717254/
  2. Mayo-Wilson E, et al. (2014). Psychological and pharmacological interventions for social anxiety disorder: network meta-analysis. Lancet Psychiatry 1:368–376. https://pubmed.ncbi.nlm.nih.gov/26361000/
  3. Carpenter JK, et al. (2018). Cognitive behavioral therapy for anxiety and related disorders: meta-analysis of randomized placebo-controlled trials. Depress Anxiety 35:502–514. https://pubmed.ncbi.nlm.nih.gov/29451967/
  4. van Dis EAM, et al. (2020). Long-term outcomes of cognitive behavioral therapy for anxiety-related disorders. JAMA Psychiatry 77:265–273. https://pubmed.ncbi.nlm.nih.gov/31758858/
  5. Papola D, et al. (2024). Psychotherapies for generalized anxiety disorder in adults: network meta-analysis. JAMA Psychiatry 81:250–259. https://pubmed.ncbi.nlm.nih.gov/37851421/
  6. Craske MG, et al. (2014). Maximizing exposure therapy: an inhibitory learning approach. Behav Res Ther 58:10–23. https://pubmed.ncbi.nlm.nih.gov/24864005/
  7. Weisman JS, Rodebaugh TL (2018). Exposure therapy augmentation: a review and extension of techniques informed by an inhibitory learning approach. Clin Psychol Rev 59:41–51. https://pubmed.ncbi.nlm.nih.gov/29128146/
  8. Mataix-Cols D, et al. (2017). D-cycloserine augmentation of exposure-based CBT: individual participant data meta-analysis. JAMA Psychiatry 74:501–510. https://pubmed.ncbi.nlm.nih.gov/28122091/
  9. Slee A, et al. (2019). Pharmacological treatments for generalised anxiety disorder: network meta-analysis. Lancet 393:768–777. https://pubmed.ncbi.nlm.nih.gov/30712879/
  10. Guaiana G, et al. (2023). Pharmacological treatments in panic disorder in adults: network meta-analysis. Cochrane Database Syst Rev CD012729. https://pubmed.ncbi.nlm.nih.gov/38014714/
  11. Henssler J, et al. (2024). Incidence of antidepressant discontinuation symptoms: systematic review and meta-analysis. Lancet Psychiatry 11:526–535. https://pubmed.ncbi.nlm.nih.gov/38851198/
  12. Batelaan NM, et al. (2017). Risk of relapse after antidepressant discontinuation in anxiety disorders, OCD and PTSD. BMJ 358:j3927. https://pubmed.ncbi.nlm.nih.gov/28903922/
  13. Takeshima N, et al. (2021). Psychosocial interventions for benzodiazepine discontinuation. Psychiatry Clin Neurosci 75:119–127. https://pubmed.ncbi.nlm.nih.gov/33448517/
  14. Steenen SA, et al. (2016). Propranolol for the treatment of anxiety disorders: systematic review and meta-analysis. J Psychopharmacol 30:128–139. https://pubmed.ncbi.nlm.nih.gov/26487439/
  15. Hedman-Lagerlöf E, et al. (2023). Therapist-supported internet-based cognitive behaviour therapy versus face-to-face CBT. World Psychiatry 22:305–314. https://pubmed.ncbi.nlm.nih.gov/37159350/
  16. Linardon J, et al. (2024). Current evidence on the efficacy of mental health smartphone apps. World Psychiatry 23:139–149. https://pubmed.ncbi.nlm.nih.gov/38214614/
  17. Linardon J, et al. (2024). The efficacy of mindfulness apps on symptoms of depression and anxiety. Clin Psychol Rev 107:102370. https://pubmed.ncbi.nlm.nih.gov/38056219/
  18. Stubbs B, et al. (2017). An examination of the anxiolytic effects of exercise for people with anxiety and stress-related disorders. Psychiatry Res 249:102–108. https://pubmed.ncbi.nlm.nih.gov/28088704/
  19. Goyal M, et al. (2014). Meditation programs for psychological stress and well-being. JAMA Intern Med 174:357–368. https://pubmed.ncbi.nlm.nih.gov/24395196/
  20. Hoge EA, et al. (2023). Mindfulness-based stress reduction vs escitalopram for anxiety disorders. JAMA Psychiatry 80:13–21. https://pubmed.ncbi.nlm.nih.gov/36350591/
  21. Hoge EA, et al. (2025). Videoconference-delivered MBSR versus escitalopram. J Affect Disord 384:163–172. https://pubmed.ncbi.nlm.nih.gov/40324655/
  22. Ferreira MG, et al. (2022). Acceptance and commitment therapy for anxiety and depression: meta-analysis. J Affect Disord 309:297–308. https://pubmed.ncbi.nlm.nih.gov/35489560/
  23. Balban MY, et al. (2023). Brief structured respiration practices enhance mood and reduce physiological arousal. Cell Rep Med 4:100895. https://pubmed.ncbi.nlm.nih.gov/36630953/
  24. Wilson J, et al. (2026). Cannabinoids for mental disorders and substance use disorders: systematic review and meta-analysis. Lancet Psychiatry 13:304–315. https://pubmed.ncbi.nlm.nih.gov/41856154/
  25. Pittler MH, Ernst E (2003). Kava extract for treating anxiety. Cochrane Database Syst Rev CD003383. https://pubmed.ncbi.nlm.nih.gov/12535473/
  26. Kasper S, et al. (2014). Efficacy of orally administered Silexan in patients with anxiety-related restlessness. Int J Neuropsychopharmacol 17:859–869. https://pubmed.ncbi.nlm.nih.gov/24456909/
  27. Turner EH, et al. (2008). Selective publication of antidepressant trials and its influence on apparent efficacy. NEJM 358:252–260. https://pubmed.ncbi.nlm.nih.gov/18199864/
  28. Ross S, et al. (2016). Rapid and sustained symptom reduction following psilocybin treatment in patients with life-threatening cancer. J Psychopharmacol 30:1165–1180. https://pubmed.ncbi.nlm.nih.gov/27909164/
  29. Griffiths RR, et al. (2016). Psilocybin produces substantial and sustained decreases in depression and anxiety in life-threatening cancer. J Psychopharmacol 30:1181–1197. https://pubmed.ncbi.nlm.nih.gov/27909165/
  30. Bandelow B, et al. (2014). Efficacy of treatments for anxiety disorders: a meta-analysis. Dtsch Arztebl Int 111:473–480. https://pubmed.ncbi.nlm.nih.gov/25138725/
  31. NICE CG113. Generalised anxiety disorder and panic disorder in adults: management. https://www.nice.org.uk/guidance/cg113

Appendix C — Register of unverified, corrected and miscited claims

Listed so that nobody launders them onward. This register is long by design.

Corrected during verification

Claim as first draftedCorrection
Reduced total choline is the only replicated MRS finding in anxietyNAA was also reduced across all cortical regions. tCho is the only finding significant with and without outlier exclusion.
ENIGMA panic effects are all d = −0.07 to −0.13True for case-control contrasts, but omits the paper's largest effect: early-onset panic disorder → larger lateral ventricles, d = 0.31–0.38.
CBD shows no effect on anxiety across 54 trialsThe null is reported for cannabinoids generally; the abstract does not isolate a CBD-only anxiety stratum, and the full text is paywalled.
The ABM registered replication found Bayesian null evidence on all measuresFour of five. The depression measure was insensitive, not null-supporting.
"Anxiet-" occurs twice in Moncrieff et al. (2022)Once in readable text (two raw-XML hits resolve to the same reference). The scope conclusion is unaffected.

Miscitations found in the commissioning briefs

  • Baxter et al., "Challenging the myth of an 'epidemic'" is in Depression and Anxiety 31:506–516 — not Psychological Medicine. (A different 2014 Baxter paper is in Psychol Med.)
  • Schmidt, Zvolensky & Maner (2006) is in Journal of Psychiatric Research 40:691–699 — not Journal of Abnormal Psychology.
  • Hancock & Ganey (2003) is in Journal of Human Performance in Extreme Environments 7:5–14 — not Journal of Performance Psychology.
  • Costello et al. (2019) is in BMJ Open — not Translational Psychiatry.
  • Mogg & Bradley's substantive reappraisal is 2016, Behaviour Research and Therapy 87:76–108; the 2018 paper is Trends in Cognitive Sciences.

Could not be verified against a primary source

ItemStatus
MacLeod, Mathews & Tata (1986) — N, design, effect sizeNo abstract indexed anywhere. The single most-cited experiment in the field is unverifiable from open sources.
Schmukle (2005) — exact reliability coefficientsClosed access. Only the verbatim "completely unreliable" conclusion is verified.
Rodebaugh et al. (2016) — specific reliability coefficientsPMC and publisher both blocked.
Duits et al. (2015) — ALL effect sizes and CIsClosed access; no citing OA source quotes them. Only directions and sample counts verified. Anyone citing a specific d should be asked to show the table.
Browning et al. (2015) — N and effect sizeNot in the abstract.
von der Embse et al. (2018) — every effect sizeThe abstract contains no r, k-per-outcome, N, CI, heterogeneity or bias statistic. Only "238 studies" is verifiable.
Etkin & Wager (2007) — study count; reports no effect sizes or CIs at allCoordinate-based meta-analysis.
Chavanne & Robinson (2021) — whether the amygdala appeared in the pathological-anxiety-only mapFull text not obtained.
Renna et al. (2018) — CIs, I², publication-bias testsPaywalled. Do not quote CIs.
Stalder (2017) hair-cortisol "−17%" in anxietyNo CI, k or p reported.
Balban et al. (2023) — analysed N, anxiety-scale effect size, attritionPMC blocked; abstract reports p-values only.
HRV biofeedback / slow-paced breathing for anxietyNo primary meta-analysis meeting source standards located. No tier assigned.
Quantitative benzodiazepine dependence risk per person-yearNo high-quality prospective estimate located. A genuine evidence gap.
Inhibitory-learning exposure vs standard exposureNo adequately powered clinical RCT establishing superiority located.
SSRI dose-response in anxiety disorders specificallyNo dedicated meta-analysis retrieved.
Combination psychotherapy + pharmacotherapy vs monotherapyCochrane found inadequate evidence; no superseding synthesis.
Stubbs et al. (2017) exercise CIThe published interval is arithmetically impossible (point estimate lies outside it). True bounds unrecoverable from the abstract.
Baxter et al. (2014, Psychol Med) UI for 390 DALYs/100,000Printed as "191–371," which excludes its own point estimate. Apparent typo, unresolved.
Kendler et al. (1992) exact genetic correlationOnly the qualitative "completely shared" wording verified.
Any direct DIF / measurement-invariance study on core diagnostic instruments (CIDI modules) by sexSearched extensively; not located. A real gap.
Any published empirical test of the prevalence inflation hypothesisNot located. Three years after the call to test it.
Any meta-analysis of risk-taking or ambiguity aversion in anxiety disordersNot located.
Any preregistered direct replication of the foundational choking studiesNot located.
Any well-powered study pitting avoidance against arousal as the driver of impairmentNot located. Theoretically central, empirically untested.
Ketamine for diagnosed anxiety disordersNo dedicated meta-analysis exists.
L-theanine for anxietyOnly the acute cognitive outcome was verifiable.

Widely quoted figures whose provenance fails

FigureProblem
"US$1 trillion per year" (depression + anxiety)Traced to a 2016 WHO press release. Not computable from its usual citation, which reports 15-year cumulative net present values ($230bn depression + $169bn anxiety productivity gain over 2016–2030). No arithmetic path to $1tn/year.
"$6 trillion by 2030"Grey literature (WEF/Harvard, 2011), and covers all mental illness, not anxiety and depression.
WHO's "359 million with an anxiety disorder in 2021"Not new primary epidemiology — it propagates the modelled COVID adjustment forward from GBD 2019's 301.4 million. Citing it alongside the null same-cohort COVID finding is citing two mutually inconsistent claims.
"1 in 9 → 1 in 5" UK youth mental health trendStraddles an instrument change (DAWBA diagnostic interview → SDQ screening questionnaire). Do not cite.
GAD's 6-month criterion "decreased to 3 months in DSM-5"Incorrect. Describes a rejected pre-publication draft proposal.

Preprints and non-peer-reviewed sources used

  • Satti et al. (2025), volatility learning, N = 820 — bioRxiv preprint, no journal version.

Things that do not exist — each reportable as a finding

A meta-analysis of the cortisol awakening response in anxiety disorders · a meta-analysis of basal or diurnal cortisol in GAD, panic or social anxiety · a dexamethasone-suppression or dex/CRH meta-analysis in anxiety · a meta-analysis of benzodiazepine-receptor imaging in anxiety · any benzodiazepine-receptor PET/SPECT study in anxiety published 2009–2026 · any neuromelanin-MRI study of the locus coeruleus in any anxiety disorder · a Cochrane review of probiotics for anxiety · a meta-analysis of oestradiol × fear extinction · an anxiety-specific equivalent of Border et al. (2019) · any CO₂ challenge ever performed in congenital central hypoventilation syndrome patients, 32 years after the theory was proposed.


Prepared 27 July 2026. Four domain dossiers (neurobiology; cognition and function; taxonomy and epidemiology; evidence-ranked interventions) and one adversarial verification pass covering 17 load-bearing claims underlie this synthesis. Companion to Stress, Energy and the Capacity to Function (26 July 2026); §2.12 and §5.2 draw the cross-review comparison.

About the author

Paul Stephen

Founder, Apatheia Labs

Evidence-governed research publication — Prosoche applied in the open.

All essays

Follow

New work, in your reader

New essays and audits publish to RSS — no inbox, no list. Point your reader at the feed and they arrive as they land.

Subscribe via RSS

Published by Apatheia Labs. All rights reserved. Quote freely with attribution; redistribute with permission.