Evidence Review
Depression
Antidepressants, the serotonin story, and what the trials actually support
On this page20 sections
Companion reviews: Antipsychotics and the Evidence, Anxiety — a critical review of the evidence, and Stress, Energy and the Capacity to Function.
Executive summary
Two folk models dominate public discussion of depression: the chemical-imbalance story (depression is low serotonin; antidepressants fix it) and its inversion (the serotonin hypothesis is wrong; therefore antidepressants don't work). Both are overclaims, and the evidence supports neither.
Eight things the evidence actually establishes:
-
Antidepressants work modestly better than placebo, but publication bias inflated their apparent efficacy by 32%. Across 74 FDA-registered trials, the published literature suggested 94% were positive; the FDA data showed 51% (Turner et al., 2008). The largest network meta-analysis of 522 trials found all 21 antidepressants more effective than placebo, with effect sizes (odds ratios) ranging from 1.37 to 2.13, but most comparative analyses carried wide credible intervals. Of those trials, 9% were rated as high risk of bias, 73% as moderate, and 18% as low; certainty of evidence ranged from moderate to very low (Cipriani et al., 2018).
-
The drug–placebo difference is severity-dependent and driven by reduced placebo response in severe depression, not increased drug response. At moderate depression severity, the difference is clinically insignificant. It reaches conventional criteria for clinical significance only in patients with baseline Hamilton Depression Rating Scale scores above 28—the upper end of very severe depression. The mechanism is not that medication works better in severe cases, but that placebo works worse (Kirsch et al., 2008).
-
First-line remission rates are 37–46%, meaning more than half of patients do not achieve full symptom resolution with their initial treatment. STAR*D found remission rates of 36.8%, 30.6%, 13.7% and 13.0% across four successive treatment steps, with a cumulative remission rate of 67% (Rush et al., 2006). Psychotherapy and pharmacotherapy produce statistically indistinguishable remission rates of approximately 46% each in direct comparisons; control conditions produce 24% (Casacalenda et al., 2002).
-
The serotonin-deficiency hypothesis of depression lacks consistent empirical support. An umbrella review of 17 studies found no convincing evidence that depression is associated with lowered serotonin concentrations or activity (Moncrieff et al., 2022). Two meta-analyses of 5-HIAA (the serotonin metabolite) in over 1,000 participants showed no association with depression. Cohort studies of plasma serotonin showed no relationship with depression, and evidence that lowered serotonin was associated with antidepressant use.
-
The serotonin hypothesis critique has been criticised for methodological flaws, selective reporting and misinterpretation of neuropsychopharmacological findings. [contested] Thirty-five co-authors argued that Moncrieff et al. (2022) mischaracterised 5-HT1A receptors as exclusively presynaptic autoreceptors—reversing a key directional inference—and ignored tryptophan depletion effects in at-risk populations (Jauhar et al., 2023). The umbrella review's conclusion is stronger than its evidence base supports, but its methodological critique—that medication confounding, directional ambiguity and underpowering plague the serotonin literature—transfers cleanly.
-
Continuing antidepressants after remission reduces relapse risk from 40% to 20% over 6–12 months. Meta-analysis of 40 studies (8,890 participants) found an odds ratio of 0.38 for relapse with continuation versus placebo (Sim et al., 2021). The treatment effect persists for up to 36 months, though most trials are of 12 months' duration. Critically, this benefit comes from enrichment-design trials that exclude non-responders before randomisation, meaning these figures apply to people who already responded to the medication.
-
One in three patients experiences at least one discontinuation symptom after stopping antidepressants. Incidence of discontinuation symptoms is 31% (95% CI 27–35%) after antidepressants versus 17% (14–21%) after placebo. The summary difference between antidepressant and placebo groups in RCTs is 8% (4–12%). Considering non-specific effects, the incidence of antidepressant discontinuation symptoms is approximately 15% (Henssler et al., 2024). Mean symptom severity at one week falls below the threshold for clinically significant discontinuation syndrome (Kalfas et al., 2025). Dizziness, nausea, vertigo and nervousness are the most common symptoms. Venlafaxine, desvenlafaxine, imipramine, paroxetine and escitalopram show higher rates of discontinuation symptoms.
-
The inference from drug efficacy to pathophysiology is a logical fallacy. Aspirin relieves headaches, but headaches are not caused by aspirin deficiency. SSRIs modestly reduce depressive symptoms, but this does not prove depression results from serotonin deficiency. The ex juvantibus inference—that a treatment's efficacy validates the disease model it was designed to address—is invalid (Lacasse & Leo, 2005).
What the evidence does not establish:
- That depression is a serotonin-deficiency disorder
- That antidepressants "correct a chemical imbalance"
- That antidepressants are ineffective (they are effective, but modestly)
- That the drug–placebo difference is large or consistent across severity levels
- Which patients will respond to which treatments
- Whether longer-term antidepressant use (beyond 12–36 months) continues to provide benefit
The structural problem that conditions all of this: Publication bias is severe, measurement is almost entirely self-report in unblinded participants, and the field's most influential efficacy claims rest on waitlist comparators that inflate effect sizes. The Turner (2008) analysis is unambiguous: among 74 FDA-registered trials, 31% were never published, and trials with negative or questionable results were either withheld or published as positive. The 32% effect-size inflation is not a hypothesis—it is a measured, replicated finding in the regulatory data.
How to read this review
The single most important heuristic
An effect size without a named comparator is not a finding. Antidepressants show larger effects against waitlist controls than against pill placebo. Psychotherapy shows larger effects against waitlist than against psychological placebo or treatment-as-usual. Any effect size quoted without its comparator should be assumed to be the most inflated version available, and mentally adjusted downward.
Three structural problems
Publication bias. Turner et al. (2008) obtained FDA reviews for 12 antidepressant agents (74 studies, 12,564 participants). Among studies the FDA deemed positive, 37 of 38 were published. Among studies the FDA deemed negative or questionable, only 3 of 36 were published as negative; 22 were not published and 11 were published in a way that conveyed a positive outcome. The published literature suggested 94% of trials were positive; the FDA data showed 51%. Effect-size inflation ranged from 11% to 69% for individual drugs and was 32% overall.
Enrichment design. Most relapse-prevention trials use an enrichment design: patients who do not respond or cannot tolerate the medication during an open-label phase are excluded before randomisation. This means relapse rates (20% on medication versus 40% on placebo) apply only to the subset of patients who already tolerated and responded to the drug. The figures do not generalise to all patients starting treatment.
Severity thresholds. Kirsch et al. (2008) analysed FDA data for four new-generation antidepressants (fluoxetine, venlafaxine, nefazodone, paroxetine). Drug–placebo differences increased as a function of baseline severity, but reached conventional criteria for clinical significance (3 points on the Hamilton Depression Rating Scale) only for patients with baseline HRSD scores above 28. At moderate severity, the difference was clinically insignificant. The relationship was attributable to decreased placebo response in very severe depression, not increased drug response.
Conventions
Claims are marked [contested] where the literature genuinely disagrees and [unverified] where a figure is widely repeated but could not be traced to a primary source in preparing this review.
Confidence intervals are omitted where they could not be verified against the primary text, rather than reconstructed.
Part I — The serotonin hypothesis: what the evidence does and does not support
1.1 The umbrella review, and what it actually found
Moncrieff et al. (2022) conducted a systematic umbrella review of serotonin and depression, searching six research areas: serotonin and 5-HIAA concentrations in body fluids, 5-HT1A receptor binding, serotonin transporter levels, tryptophan depletion studies, SERT gene associations, and SERT gene–environment interactions. The review included 17 studies: 12 systematic reviews and meta-analyses, one collaborative meta-analysis, one meta-analysis of large cohort studies, one systematic review and narrative synthesis, one genetic association study and one umbrella review.
The main findings:
| Domain | Finding | Source |
|---|---|---|
| 5-HIAA (serotonin metabolite) | Two meta-analyses of overlapping studies (largest n = 1,002): no association with depression | Moncrieff et al., 2022 |
| Plasma serotonin | One meta-analysis of cohort studies (n = 1,869): no relationship with depression; lowered serotonin associated with antidepressant use | Moncrieff et al., 2022 |
| 5-HT1A receptor binding | Two meta-analyses of overlapping studies (largest n = 561): weak and inconsistent evidence of reduced binding in some areas; would be consistent with increased synaptic serotonin if causal | Moncrieff et al., 2022 |
| SERT binding | Three meta-analyses of overlapping studies (largest n = 1,845): weak and inconsistent evidence; effects of prior antidepressant use not reliably excluded | Moncrieff et al., 2022 |
| Tryptophan depletion | One meta-analysis (n = 566 healthy volunteers): no effect in most; weak evidence of effect in those with family history of depression (n = 75) | Moncrieff et al., 2022 |
| SERT gene (5-HTTLPR) | Two largest studies (n = 115,257 and n = 43,165): no association | Moncrieff et al., 2022 |
The authors concluded: "there is no convincing evidence that depression is associated with, or caused by, lower serotonin concentrations or activity."
What this conclusion supports: The popular "chemical imbalance" model—that depression results from low serotonin in the way that diabetes results from low insulin—lacks empirical support. The evidence base is weak, inconsistent and confounded.
What it does not support: That serotonin plays no role in depression, or that antidepressants are ineffective. The umbrella review addresses aetiology (whether low serotonin causes depression), not treatment efficacy (whether drugs that act on serotonin reduce symptoms).
1.2 The rebuttal, and the fallacy both sides commit
Jauhar et al. (2023), signed by 35 researchers, argued that Moncrieff et al. (2022) was "methodologically flawed" and "inconsistent with the conventional umbrella review process." They identified several problems:
-
Directional mischaracterisation: Moncrieff et al. described reduced 5-HT1A receptor binding as consistent with increased serotonin, treating 5-HT1A receptors as exclusively presynaptic autoreceptors. But postsynaptic 5-HT1A receptors exist, and reduced binding at postsynaptic sites would indicate decreased serotonergic activity, not increased. This reverses a key inference.
-
Selective omission: The tryptophan depletion literature shows that acute tryptophan depletion does not provoke depressive symptoms in healthy volunteers, but does in people with a history of depression or a family history of depression. Moncrieff et al. reported the null in healthy volunteers but gave less weight to the positive finding in at-risk groups.
-
Umbrella methodology: Unlike most umbrella reviews, Moncrieff et al. summarised existing results without extracting data or conducting new analysis. Jauhar et al. argue this is inconsistent with the umbrella review method.
[contested] The rebuttal is substantive. But note the boomerang: after criticising Moncrieff for discussing efficacy in a review of aetiology, Jauhar et al. close with "The proven efficacy of SSRIs… lends credibility to this position." This is the ex juvantibus fallacy—the inference that a treatment's efficacy validates the disease model it was designed to address—in the opposite direction.
Lacasse & Leo (2005) demolished this inference: "the fact that aspirin cures headaches does not prove that headaches are due to low levels of aspirin in the brain." Antidepressants' modest efficacy tells us they reduce symptoms; it does not tell us that low serotonin caused the symptoms.
What survives: The serotonin hypothesis as a folk model ("depression is a chemical imbalance") is not supported by consistent evidence. The serotonin system as one component in a complex, multi-system disorder remains plausible. The mechanistic critique—that biomarker studies are small, confounded by medication, and directionally ambiguous—is sound and applies to both sides of this debate.
1.3 The medication confound, and why it matters
Several of the findings Moncrieff et al. cite as evidence against the serotonin hypothesis may reflect the effects of antidepressant treatment, not the state of depression itself.
Plasma serotonin: In cohort studies, lowered plasma serotonin was associated with antidepressant use, not with depression (Moncrieff et al., 2022). This is consistent with the hypothesis that long-term SSRI use causes compensatory downregulation.
SERT binding: Imaging studies of SERT availability show "weak and inconsistent evidence," and effects of prior antidepressant use were not reliably excluded (Moncrieff et al., 2022). Antidepressants directly target SERT, so any population with high rates of past or current medication use will show altered SERT binding independent of the disorder.
5-HT1A receptor binding: Similarly, reduced 5-HT1A binding could reflect chronic SSRI exposure rather than pre-existing pathology.
The field's central problem is that the population in which we measure biomarkers (people diagnosed with depression) has high exposure to drugs that directly perturb the system we are trying to measure. Disentangling cause from treatment effect requires large, medication-naive cohorts followed prospectively—and these studies are rare.
The directional ambiguity is worse than Moncrieff acknowledges. Reduced postsynaptic receptor binding can mean either (1) chronic overstimulation leading to downregulation, or (2) understimulation and reduced need for receptors. Without independent measures of synaptic serotonin concentration, receptor data alone cannot distinguish the two.
Part II — Efficacy: what antidepressants actually do, and for whom
2.1 The FDA data, and the 32% inflation
Turner et al. (2008) obtained FDA reviews for studies of 12 antidepressant agents involving 12,564 patients. They conducted a systematic literature search to identify matching publications, compared published outcomes with FDA outcomes, and compared the effect size derived from published reports with the effect size derived from the entire FDA dataset.
The findings:
| Outcome | FDA data | Published literature |
|---|---|---|
| Number of studies | 74 | 51 published |
| Studies deemed positive by FDA | 38 (51%) | 37 published |
| Studies deemed negative/questionable by FDA | 36 (49%) | 3 published as negative; 11 published as positive; 22 not published |
| Apparent success rate | 51% | 94% |
| Effect-size inflation | — | 11–69% per drug; 32% overall |
The magnitude is unambiguous. Studies viewed by the FDA as positive were approximately 12 times as likely to be published in agreement with the FDA analysis as studies with non-positive results (risk ratio 11.7, 95% CI 6.2–22.0, p < 0.001).
Turner et al. conclude: "We cannot determine whether the bias observed resulted from a failure to submit manuscripts on the part of authors and sponsors, from decisions by journal editors and reviewers not to publish, or both."
This is not a historical problem. The Turner analysis covers trials from 1987 to 2004, but the Cipriani et al. (2018) network meta-analysis—covering 522 trials up to 2016—still found that 9% of trials were rated as high risk of bias and 73% as moderate risk, with certainty of evidence ranging from moderate to very low. Preregistration has improved the situation for drugs, but the field's foundational efficacy estimates were built on a selectively published evidence base.
2.2 The severity interaction: placebo falls, drugs stay flat
Kirsch et al. (2008) obtained FDA data for four new-generation antidepressants (fluoxetine, venlafaxine, nefazodone, paroxetine) and used meta-analytic techniques to assess the relationship between baseline severity and improvement.
Key findings:
- Drug–placebo differences increased as a function of baseline severity, rising from virtually no difference at moderate levels to a relatively small difference at very severe levels.
- Clinical significance (defined as a 3-point difference on the 17-item Hamilton Depression Rating Scale) was reached only for patients with baseline HRSD scores above 28—the upper end of the very severely depressed category.
- Meta-regression showed the relationship was curvilinear in drug groups (slight increase with severity) and showed a strong negative linear component in placebo groups (placebo response falls sharply with increasing severity).
The mechanism is not that medication works better in severe depression. It is that placebo works worse. At moderate severity, both medication and placebo produce similar improvement. At very high severity, placebo improvement drops, widening the gap.
Clinical interpretation: For patients with mild-to-moderate depression (HRSD < 23), the drug–placebo difference may not be clinically meaningful. For very severe depression (HRSD > 28), antidepressants provide a clinically significant benefit over placebo. The threshold matters.
[contested] This analysis has been criticised on the grounds that the four drugs studied (especially nefazodone, which was withdrawn) are not representative of the full class. The Cipriani et al. (2018) analysis included 21 antidepressants and did not replicate the same severity-by-treatment interaction. But Cipriani pooled across all severity levels without testing for moderation, so the two analyses address different questions.
2.3 The network meta-analysis: all drugs beat placebo, but not by much
Cipriani et al. (2018) conducted a systematic review and network meta-analysis of 522 double-blind randomised controlled trials (116,477 participants) comparing 21 antidepressants for acute treatment of major depressive disorder in adults.
Efficacy (response rate, defined as ≥50% symptom reduction):
All 21 antidepressants were more effective than placebo. Odds ratios ranged from 2.13 (95% CrI 1.89–2.41) for amitriptyline to 1.37 (1.16–1.63) for reboxetine.
When all trials were considered, differences in odds ratios between antidepressants ranged from 1.15 to 1.55 for efficacy, with wide credible intervals on most comparative analyses.
In head-to-head studies, agomelatine, amitriptyline, escitalopram, mirtazapine, paroxetine, venlafaxine and vortioxetine were more effective than other antidepressants (ORs 1.19–1.96). Fluoxetine, fluvoxamine, reboxetine and trazodone were the least efficacious (ORs 0.51–0.84).
Acceptability (treatment discontinuation due to any cause):
Only agomelatine (OR 0.84, 95% CrI 0.72–0.97) and fluoxetine (0.88, 0.80–0.96) were associated with fewer dropouts than placebo. Clomipramine was worse than placebo (1.30, 1.01–1.68).
Quality of evidence: 9% of 522 trials (46 trials) were rated as high risk of bias, 73% (380 trials) as moderate, and 18% (96 trials) as low. Certainty of evidence ranged from moderate to very low.
The authors' conclusion: "All antidepressants were more efficacious than placebo in adults with major depressive disorder. Smaller differences between active drugs were found when placebo-controlled trials were included in the analysis, whereas there was more variability in efficacy and acceptability in head-to-head trials."
Translation: Antidepressants work. They work better than placebo. The differences between drugs are smaller than marketing suggests, and most comparative evidence is of moderate to low quality.
2.4 Remission rates: more than half do not achieve full symptom resolution
Direct comparisons (Casacalenda et al., 2002): A review of six multiple-cell randomised controlled trials (883 outpatients with mild-to-moderate, primarily non-melancholic, non-psychotic major depressive disorder) found remission percentages of 46.4% for medication, 46.3% for psychotherapy, and 24.4% for control conditions. Treatment duration ranged from 10 to 34 weeks (median 16 weeks). Dropout rates were 37.1% for medication, 22.2% for psychotherapy, and 54.4% for control conditions.
STAR*D (Rush et al., 2006): The Sequenced Treatment Alternatives to Relieve Depression trial followed 3,671 outpatients with non-psychotic major depressive disorder through up to four successive treatment steps. Remission rates were 36.8% for step one, 30.6% for step two, 13.7% for step three, and 13.0% for step four. The overall cumulative remission rate was 67%.
Critically, patients who required more treatment steps had higher relapse rates during the 12-month naturalistic follow-up phase. Lower relapse rates were found among participants who were in remission at follow-up entry than for those who were not.
Meta-analysis of psychotherapy (Cuijpers et al., 2022): Across 441 trials, psychotherapy response rates (≥50% symptom reduction) were 42% for major depressive disorder, 38% for PTSD, 38% for OCD, 38% for panic disorder, 36% for GAD, and 32% for social anxiety disorder. Control conditions (waitlist, care-as-usual, pill placebo) produced response rates of 17–24%.
What these numbers mean: First-line treatment—whether medication or psychotherapy—produces remission or response in roughly half of patients. The remainder require additional steps, and remission rates drop sharply in later steps. By the time a patient reaches a third or fourth treatment attempt, the probability of remission is 13–14%.
2.5 Relapse after discontinuation: the enrichment problem
The headline finding (Sim et al., 2021): Meta-analysis of 40 studies (8,890 participants) found that continuing antidepressants after remission reduced relapse rates from 39.7% on placebo to 20.9% on antidepressant over periods ranging from 14 to 100 weeks (median ~40 weeks). Odds ratio: 0.38 (95% CI 0.33–0.43).
The effect was significant for TCAs (OR 0.30), SSRIs (OR 0.33), and other newer agents (OR 0.44). Even in studies where continuous treatment exceeded 6 months before randomisation, continued use had lower relapse than placebo (OR 0.40).
The critical design issue: Most of these trials use an enrichment design. Patients who do not respond to or cannot tolerate the antidepressant during an open-label run-in phase are excluded before randomisation to continuation or placebo. This means the 20% versus 40% relapse figures apply only to the subset of patients who already responded to and tolerated the medication.
Translation: If you responded well to an antidepressant, continuing it for at least 6–12 months after remission roughly halves your relapse risk compared to stopping. If you did not respond, or could not tolerate it, these figures do not apply to you—because people like you were excluded from the trials before randomisation.
How long to continue? Geddes et al. (2003) found that the two-thirds reduction in relapse risk was largely independent of the duration of treatment before randomisation or the duration of randomly allocated therapy. The effect persisted for up to 36 months, though most trials were of 12 months' duration. A more recent meta-regression (Kishi et al., 2023) found similar relapse rates at 6 months and 12 months, suggesting that continuation for at least 6 months is supported, but the marginal benefit beyond 6–12 months is unclear.
Part III — Discontinuation: withdrawal symptoms and nocebo effects
3.1 Incidence and severity
Henssler et al. (2024): Systematic review and meta-analysis of 79 studies (21,002 patients; 44 RCTs and 35 observational studies). Incidence of at least one discontinuation symptom was 31% (95% CI 27–35%) after discontinuing antidepressants and 17% (14–21%) after discontinuing placebo. Between antidepressant and placebo groups in RCTs, the summary difference in incidence was 8% (4–12%). Considering non-specific effects as evidenced in placebo groups, the authors interpret the incidence of antidepressant discontinuation symptoms as approximately 15%.
Incidence of severe discontinuation symptoms was 2.8% (1.4–5.7%) after antidepressants versus 0.6% (0.2–1.3%) after placebo.
Desvenlafaxine, venlafaxine, imipramine and escitalopram were associated with higher frequencies of discontinuation symptoms. Imipramine, paroxetine, and desvenlafaxine/venlafaxine were associated with higher severity.
Kalfas et al. (2025): Meta-analysis of 49 RCTs (17,828 participants). Participants who stopped antidepressants experienced a mean of one more discontinuation symptom compared to those who discontinued placebo or continued antidepressants (standardised mean difference 0.31, 95% CI 0.23–0.39). The effect size was equivalent to one additional symptom on the Discontinuation-Emergent Signs and Symptoms scale.
The most common symptoms were dizziness (risk difference 6.24%), nausea, vertigo, and nervousness. Mean symptom severity at one week was below the threshold for clinically significant discontinuation syndrome.
Critically, discontinuation was not associated with depressive symptoms in the first two weeks (k = 5 studies in people with major depressive disorder). This means that mood worsening in the first two weeks after stopping is more likely to reflect nocebo or expectation effects; later presentation of depression (beyond two weeks) is more consistent with relapse.
Translation: About one in three people stopping antidepressants will experience at least one discontinuation symptom. The RCT summary difference from placebo is 8%, and considering non-specific effects, the drug-attributable incidence is approximately 15%. Most symptoms are mild and resolve within one to two weeks. Severe discontinuation syndromes occur in approximately 3% of discontinuations. The symptoms are real, but they are neither universal nor typically severe.
3.2 Distinguishing discontinuation symptoms from relapse
The Kalfas et al. (2025) finding—that discontinuation is not associated with depressive symptoms in the first two weeks—provides a useful timeline:
- Days 1–14: Physical symptoms (dizziness, nausea, vertigo) are most common. Mood worsening during this window is not a typical discontinuation symptom and may reflect nocebo effects or contextual factors.
- Weeks 3+: Emergence of depressive symptoms after the discontinuation window is more consistent with relapse than with withdrawal.
This distinction matters clinically. If someone's mood deteriorates two days after stopping an antidepressant, the evidence suggests this is not a medication-withdrawal effect—it is either a nocebo response or the early return of depressive symptoms. If mood deteriorates six weeks after stopping, that is consistent with relapse.
Part IV — What the evidence does not establish
4.1 Which patients will respond to which treatments
STAR*D, the largest and most ambitious trial of sequential treatment strategies, found that 67% of patients achieved remission across up to four treatment steps. But it provided no validated algorithm for matching patients to treatments. Baseline characteristics, symptom profiles, and prior treatment history did not reliably predict which patients would respond to which medications.
The Cipriani et al. (2018) network meta-analysis found that differences in efficacy between antidepressants exist, but credible intervals were wide and overlapping for most comparisons. In clinical practice, this means trial-and-error remains the dominant treatment selection strategy.
What this means: We know that antidepressants work for some people, but we cannot currently predict who will respond or which drug will work. First-line treatment remains an empirical trial.
4.2 Whether longer-term use (beyond 12–36 months) continues to provide benefit
Relapse-prevention trials typically run for 6–18 months, with the longest extending to 36 months. Geddes et al. (2003) found that the relapse-reduction benefit persisted for up to 36 months, but stated: "the evidence on longer-term treatment requires confirmation."
There are no large, high-quality randomised trials of maintenance treatment extending beyond three years. Observational data suggest that some patients remain well on long-term maintenance, but these cohorts are subject to selection bias (people who tolerate and benefit from medication are more likely to continue it).
The key unanswered question: Does the relapse-reduction benefit plateau at some point, or does it continue indefinitely? Current evidence supports continuation for at least 6–12 months after remission, and possibly up to 36 months, but beyond that the data are sparse.
4.3 The mechanism by which antidepressants work
The monoamine hypothesis—that antidepressants work by increasing synaptic serotonin and/or norepinephrine—faces a temporal problem: SSRIs block serotonin reuptake within hours, but clinical improvement takes 2–6 weeks. If the mechanism were simply "more synaptic serotonin," the timeline does not fit.
Alternative hypotheses include:
-
Neuroplasticity: Chronic SSRI administration increases brain-derived neurotrophic factor (BDNF), promotes hippocampal neurogenesis (in animal models), and may reverse stress-induced dendritic atrophy. But whether neurogenesis is necessary or sufficient for antidepressant response in humans remains unresolved.
-
Downregulation of monoamine autoreceptors: Chronic SSRI use desensitises presynaptic 5-HT1A autoreceptors, increasing serotonin release. This matches the 2–6 week timeline better than the acute reuptake blockade.
-
Network-level effects: Antidepressants may work not by correcting a single neurotransmitter deficit, but by altering large-scale brain network dynamics. Resting-state fMRI studies show that SSRIs reduce connectivity within the default mode network and increase connectivity between the default mode network and task-positive networks. But these are correlational findings in small samples.
The honest position: We do not know the mechanism by which antidepressants reduce depressive symptoms. The serotonin hypothesis as a simple deficit model is not supported, but serotonin clearly plays some role (otherwise SSRIs would not work). The mechanism is likely multi-system, involving monoamines, neuroplasticity, inflammation and network-level changes, and the relevant mechanism may differ across patients.
Part V — The two overclaims, and what replaces them
5.1 The chemical-imbalance story
The claim: Depression is caused by low serotonin in the brain, analogous to diabetes being caused by low insulin. Antidepressants work by correcting this chemical imbalance.
What the evidence shows: There is no consistent evidence that people with depression have lower serotonin concentrations or activity than people without depression. Plasma serotonin, cerebrospinal fluid 5-HIAA, 5-HT1A receptor binding, and SERT binding studies are weak, inconsistent, and confounded by medication exposure. The two largest genetic studies of the serotonin transporter gene (5-HTTLPR) found no association with depression.
The inference from efficacy to aetiology is invalid. Aspirin cures headaches, but headaches are not caused by aspirin deficiency. SSRIs reduce depressive symptoms, but this does not prove depression is caused by serotonin deficiency. The ex juvantibus fallacy—that a treatment's efficacy validates the disease model it was designed to address—is a logical error, not an empirical claim.
What survives: Serotonin is involved in mood regulation, and drugs that act on the serotonin system modestly reduce depressive symptoms. But "involved in" is not the same as "deficient in," and the simple chemical-imbalance model does not survive contact with the biomarker data.
5.2 The "antidepressants don't work" story
The claim: Because the serotonin hypothesis is wrong, and because publication bias inflated their efficacy, antidepressants are no better than placebo or provide only trivial benefits.
What the evidence shows: All 21 antidepressants evaluated in the Cipriani et al. (2018) network meta-analysis were significantly more effective than placebo, with odds ratios ranging from 1.37 to 2.13. Turner et al. (2008) found a 32% effect-size inflation due to publication bias, but even after correcting for this, antidepressants remained more effective than placebo. Kirsch et al. (2008) found that the drug–placebo difference was small at moderate severity and clinically significant only at very high severity—but this is still a real difference, not a null finding.
Remission rates of 37–46% for first-line treatment are modest, but they are roughly double the 24% remission rate in control conditions (Casacalenda et al., 2002). Continuing antidepressants after remission reduces relapse from 40% to 20% (Sim et al., 2021). These are not large effects, but they are not trivial.
What survives: Antidepressants work modestly better than placebo, with effect sizes in the small-to-medium range. Roughly half of patients achieve remission or response with first-line treatment. The benefits are severity-dependent: larger in very severe depression, smaller in mild-to-moderate depression. Publication bias has inflated their apparent efficacy, and the evidence base is of moderate to very low certainty, but the corrected estimates remain positive.
5.3 The version that survives
Depression is a complex, multi-system disorder. It involves monoamine systems, HPA-axis dysregulation, inflammation, neuroplasticity, network-level brain changes, and environmental, genetic, and developmental factors. No single neurotransmitter deficit explains it, and no single intervention treats it in all patients.
Antidepressants work for some people, modestly, in ways we do not fully understand. First-line treatment (medication or psychotherapy) produces remission or response in roughly half of patients. The remainder require additional treatment steps, and cumulative remission rates reach 67% after four steps (Rush et al., 2006). Continuing medication after remission reduces relapse risk, but roughly 20% of people relapse even on continued treatment.
The field's foundational efficacy claims rest on a selectively published evidence base. Turner et al. (2008) documented 32% effect-size inflation due to publication bias, and the Cipriani et al. (2018) network meta-analysis found that 9% of trials were at high risk of bias, 73% at moderate risk, and certainty of evidence ranged from moderate to very low. The benefits are real, but they are smaller than the published literature suggested.
The comparator matters. Effect sizes collapse as controls get more active. Antidepressants show larger effects against waitlist than against pill placebo. Psychotherapy shows larger effects against waitlist than against psychological placebo or treatment-as-usual. Any effect size quoted without its comparator should be assumed to be inflated.
What the evidence establishes: Antidepressants modestly reduce depressive symptoms compared to placebo, with first-line remission rates of 37–46%. The benefits are larger in severe depression and smaller in mild-to-moderate depression. Continuing treatment after remission reduces relapse. Discontinuation symptoms occur in about one in three people, with an RCT summary difference of 8% from placebo and approximately 15% when accounting for non-specific effects, and most symptoms are mild. The serotonin-deficiency model lacks empirical support, but serotonin is clearly involved. Publication bias inflated apparent efficacy by 32%, but corrected estimates remain positive.
What it does not establish: That depression is a serotonin-deficiency disorder, that antidepressants correct a chemical imbalance, which patients will respond to which treatments, or whether benefits persist beyond 12–36 months.
The structural problem: Nearly all outcomes are self-reported by unblinded participants, publication bias is severe, certainty of evidence is moderate to very low, and the field's most influential claims rest on waitlist comparators. The evidence base is real but inflated, and the honest summary is smaller and more qualified than either the "chemical imbalance" or the "antidepressants don't work" narrative.
References
Primary sources cited
Casacalenda, N., Perry, J. C., & Looper, K. (2002). "Remission in major depressive disorder: A comparison of pharmacotherapy, psychotherapy, and control conditions." American Journal of Psychiatry, 159(8), 1354–1360.
Cipriani, A. et al. (2018). "Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: A systematic review and network meta-analysis." The Lancet, 391(10128), 1357–1366.
Cuijpers, P. et al. (2022). "Absolute and relative outcomes of psychotherapies for eight mental disorders: A systematic review and meta-analysis." World Psychiatry, 21, 445–451.
Geddes, J. R. et al. (2003). "Relapse prevention with antidepressant drug treatment in depressive disorders: A systematic review." The Lancet, 361(9358), 653–661.
Henssler, J. et al. (2024). "Incidence of antidepressant discontinuation symptoms: A systematic review and meta-analysis." The Lancet Psychiatry, 11(7), 526–536.
Jauhar, S. et al. (2023). "A leaky umbrella has little value: Evidence clearly indicates the serotonin system is implicated in depression." Molecular Psychiatry, 28, 3149–3152.
Kalfas, S. et al. (2025). "Incidence and nature of antidepressant discontinuation symptoms: A systematic review and meta-analysis." JAMA Psychiatry, 82(2), 135–144.
Kirsch, I. et al. (2008). "Initial severity and antidepressant benefits: A meta-analysis of data submitted to the Food and Drug Administration." PLoS Medicine, 5(2), e45.
Kishi, T. et al. (2023). "Relapse and its modifiers in major depressive disorder after antidepressant discontinuation: Meta-analysis and meta-regression." Molecular Psychiatry, 28, 3086–3096.
Lacasse, J. R., & Leo, J. (2005). "Serotonin and depression: A disconnect between the advertisements and the scientific literature." PLoS Medicine, 2(12), e392.
Moncrieff, J. et al. (2022). "The serotonin theory of depression: A systematic umbrella review of the evidence." Molecular Psychiatry, 28, 3243–3256.
Rush, A. J. et al. (2006). "Acute and longer-term outcomes in depressed outpatients requiring one or several treatment steps: A STAR*D report." American Journal of Psychiatry, 163(11), 1905–1917.
Sim, K. et al. (2021). "Discontinuation of antidepressants after remission with antidepressant medication in major depressive disorder: A systematic review and meta-analysis." Molecular Psychiatry, 26, 3346–3356.
Turner, E. H. et al. (2008). "Selective publication of antidepressant trials and its influence on apparent efficacy." New England Journal of Medicine, 358(3), 252–260.
Appendix: What this review corrected
Four claims in preliminary drafts required narrowing or correction after primary-source verification:
-
The Moncrieff et al. (2022) umbrella review scope: Early drafts stated the review "addressed depression and anxiety." Primary-source verification confirms the string "anxiet-" appears only once in the readable full text (inside a journal name in the reference list). The review concerns depression only. Corrected throughout.
-
The Turner (2008) replication: The anxiety review states that across 74 FDA-registered trials, 94% positive in published literature versus 51% in FDA data. This claim was verified against Turner et al. (2008) and is accurate for antidepressant trials registered for both depression and some anxiety indications. The figure is used here as the canonical publication-bias estimate.
-
Discontinuation symptom severity threshold: Early drafts stated "most discontinuation symptoms are mild." Kalfas et al. (2025) provide the precise threshold: mean symptom severity at one week is below the threshold for clinically significant discontinuation syndrome on the DESS. This is a more precise claim than "mild" and has been adopted.
-
The Jauhar et al. (2023) co-author count: Early drafts stated "a rebuttal signed by 34 researchers." The actual count is 35. Corrected.
About the author
Paul Stephen
Founder, Apatheia Labs
Evidence-governed research publication — Prosoche applied in the open.
Read next
More in Evidence- EssayAugust 2026
Adult ADHD and the Evidence
Prospective cohort studies find that childhood ADHD persistence to adulthood is low across named cohorts (Dunedin 5%, Pelotas 17.2%, E-Risk 21.9%), most adult-ADHD cases lack prospectively diagnosed childhood ADHD (Dunedin 90%, Pelotas 84.6%, E-Risk 67.5%), and after excluding substance use and other disorders, fewer than 1% of originally undiagnosed children meet late-onset criteria.
- EssayAugust 2026
ADHD Medications and the Evidence
Guanfacine and clonidine show child clinician SMDs −0.67 and −0.71 (Cortese 2018), but atomoxetine and viloxazine carry boxed warnings for suicidal ideation. Adult lisdexamfetamine Castells 2018 SMD −1.06; bupropion Verbeeck 2017 pooled −0.50 but clinician-only ns. Adult ER-MPH Boesen 2022 investigator SMD −0.42; IR-MPH Cândido 2021 investigator MD −20.70 but participant ns. Adult GXR Iwanami 2020 Japan diff −4.28; child GXR Sallee 2009 LS-mean −5.41 to −7.88.
- EssayAugust 2026
Antipsychotics and the Evidence
All antipsychotics beat placebo in acute trials (SMD 0.33–0.88), but 74% of CATIE participants discontinued within 18 months, dose reduction in first-episode patients doubled 7-year recovery rates versus maintenance treatment, and long-term observational data show better outcomes off medication—confounded by selection, but challenging the 'indefinite treatment for all' assumption.