Skip to content
Apatheia Labs
Writing

Evidence Review

Placebo, Expectancy, and Unblinding

What pill-placebo controls actually compare, and why recognizable side effects compromise the blind

Paul StephenApatheia LabsAugust 24, 2026 · 31 min read

Companion reviews: Antipsychotics and the Evidence, Depression — antidepressants and the evidence, ADHD — stimulants and the evidence, Anxiety — a critical review of the evidence, and Diagnostic Thresholds, Prevalence, and Overdiagnosis.


Executive summary

Two folk models dominate public discussion of psychiatric drug trials: that double-blind placebo controls cleanly isolate drug-specific effects (the orthodox position), and its inversion, that these trials measure nothing but amplified placebo effects (the radical skeptical position). Both are overclaims, and the evidence supports neither.

Seven things the evidence actually establishes:

  1. Recognizable side effects compromise blinding in pill-placebo trials. In ADHD stimulant trials, decreased appetite occurs in 266% more children on methylphenidate versus placebo (RR 3.66), and sleep problems in 60% more (RR 1.60). Among antidepressant trials, SSRI-specific side effects (sexual dysfunction, gastrointestinal disturbance) reveal treatment assignment to both participants and raters. In obesity trials using phenylpropanolamine, 74% of placebo participants and 43% of active-drug participants correctly guessed their assignment based on appetite control or lack of side effects (Moscucci et al., 1987). When guessing occurs, outcome ratings correlate with perceived assignment, independent of actual assignment.

  2. Active-placebo trials that mimic side effects show smaller drug-placebo differences than trials using inert placebos. Moncrieff et al. (2004) meta-analyzed nine trials of antidepressants versus active placebos containing atropine (to mimic anticholinergic effects). The pooled effect size was SMD 0.39 (95% CI 0.24–0.54), but this was driven by one strongly positive outlier trial. Excluding that trial reduced the pooled effect to SMD 0.17 (95% CI 0.00–0.34). For inpatient trials (predominantly severe/endogenous depression), the pooled effect was SMD 0.12 (95% CI –0.14 to 0.38). This suggests that unblinding effects inflate efficacy estimates in trials using inert placebos.

  3. The drug-placebo difference is severity-dependent and driven by reduced placebo response in severe depression, not increased drug response. Kirsch et al. (2008) analyzed FDA data for four antidepressants (fluoxetine, venlafaxine, nefazodone, paroxetine; 35 trials, 5,133 patients). Drug-placebo differences increased as a function of baseline severity, rising from virtually no difference at moderate levels of depression to a clinically significant difference only for patients with baseline Hamilton Rating Scale scores above 28—the upper end of the very severely depressed category. Meta-regression showed the relationship was curvilinear in drug groups (slight increase with severity) and showed a strong negative linear component in placebo groups (placebo response falls sharply with increasing severity). The mechanism is not that medication works better in severe depression; it is that placebo works worse.

  4. Publication bias inflated apparent antidepressant efficacy by 32%. Turner et al. (2008) obtained FDA reviews for 74 trials of 12 antidepressants (12,564 patients). Among studies the FDA deemed positive, 37 of 38 were published. Among studies the FDA deemed negative or questionable, only 3 of 36 were published as negative; 22 were not published and 11 were published in a way that conveyed a positive outcome. According to the published literature, 94% of trials were positive; the FDA data showed 51%. Effect-size inflation ranged from 11–69% for individual drugs and was 32% overall. The bias was associated with study outcome independent of sample size.

  5. Patient expectancy mediates placebo effects in antidepressant trials. In a randomized trial comparing open versus placebo-controlled citalopram (Rutherford et al., 2017), postrandomization expectancy scores were significantly higher in the open group (mean 12.1, SD 2.1) versus the placebo-controlled group (mean 11.0, SD 2.0). Hamilton Depression Rating Scale scores declined faster in the open group. Patient expectations postrandomization partially mediated group effects on week-8 scores, independent of actual drug assignment. This demonstrates that knowing one's probability of receiving active treatment—50% versus 100%—affects symptom trajectories.

  6. Nocebo effects are substantial and resemble the side-effect profile of the active drug. In antidepressant trials, placebo-arm participants report side effects at rates closely matching those in active-drug arms, including drug-specific effects. Rief et al. found that greater tolerance of SSRIs versus tricyclics was identical in both placebo and active groups, indicating that awareness of expected side effects influences tolerability independent of drug exposure. After healthy volunteers received 50 mg amitriptyline for 4 days, subsequent placebo administration reproduced amitriptyline-specific side effects, demonstrating that nocebo responses can be conditioned.

  7. Control condition type affects remission rates. Casacalenda et al. (2002) found remission rates of 46.4% for medication, 46.3% for psychotherapy, and 24.4% for control conditions in direct comparisons of outpatients with mild-to-moderate depression. Control conditions across the six included trials were pill placebo plus clinical management, treatment as usual by a general physician, and low-intensity supportive therapy. Effect sizes vary by comparator type; any effect size quoted without naming its comparator should be assumed to be the most inflated version available.

What the evidence does not establish:

  • That psychiatric drug effects are entirely placebo effects (they are not; corrected estimates remain positive)
  • That double-blind designs cleanly isolate drug-specific effects (they do not; unblinding and expectancy contaminate the estimate)
  • That placebo effects are inert or trivial (they are not; they account for most of the observed improvement in both arms)
  • That active-placebo controls are a complete solution (they reduce but do not eliminate expectancy effects, and ethical concerns prevent their widespread use)
  • Which patients will experience clinically meaningful drug-specific effects versus placebo-level effects

The structural problem that conditions all of this: Pill-placebo trials are not inert-versus-active comparisons. They compare active-drug-with-expectancy against inert-substance-with-uncertain-expectancy in a context where recognizable side effects reveal treatment assignment. The drug-placebo difference combines (1) drug-specific pharmacological effects, (2) differential expectancy effects due to unblinding, and (3) in the case of selective publication, an additional 32% inflation from reporting bias. Disentangling these components requires active placebos, better blinding strategies, and complete trial registration—none of which are standard practice.


How to read this review

The single most important heuristic

A drug-placebo difference without a named comparator and a statement about blinding integrity is not a clean estimate. Antidepressants show larger effects against inert placebo than against active placebo. Drugs and psychotherapy both show larger effects against waitlist than against pill or psychological placebo. Any efficacy claim that does not specify (1) what the control condition was, and (2) whether side effects compromised blinding, should be mentally adjusted downward.

Three structural problems

Unblinding due to recognizable side effects. In methylphenidate trials, decreased appetite affects 266% more children on drug versus placebo. In antidepressant trials, SSRI sexual dysfunction, nausea, and dry mouth are easily recognized. When patients and raters correctly guess treatment assignment, outcome ratings correlate with perceived assignment rather than actual assignment. Active placebos that mimic side effects reduce this bias but are rarely used.

Severity-dependent placebo response. The drug-placebo difference in depression is not constant across severity levels. Kirsch et al. (2008) found that the difference reached clinical significance only at HRSD scores above 28. The pattern was driven by reduced placebo response in severe depression, not increased drug response. This means that at moderate severity levels, the drug-placebo difference is clinically insignificant.

Publication bias. Turner et al. (2008) documented that 31% of FDA-registered antidepressant trials were not published, and 11 additional trials were published as positive despite FDA conclusions of negative or questionable results. The 32% effect-size inflation is not a hypothesis—it is a measured, replicated finding. Pre-registration has improved the situation for newer drugs, but the field's foundational efficacy estimates rest on a selectively published evidence base.

Conventions

Claims are marked [contested] where the literature genuinely disagrees and [unverified] where a figure is widely repeated but could not be traced to a primary source in preparing this review.

Confidence intervals are omitted where they could not be verified against the primary text, rather than reconstructed.


Part I — What pill-placebo controls actually compare

1.1 The logic of the pill-placebo design

A randomized, double-blind, placebo-controlled trial is designed to isolate the drug-specific effect by comparing outcomes in patients receiving active medication against outcomes in patients receiving an inert substance, while ensuring that neither patients nor raters know who received which treatment.

The ideal: If blinding is perfect, then the only difference between groups is the pharmacological activity of the drug. Any difference in outcomes can therefore be attributed to the drug's specific mechanism.

The reality: Blinding is compromised when patients or raters can distinguish active drug from placebo based on recognizable side effects. When this occurs, the two groups no longer differ only in pharmacology—they also differ in expectancy, attention to symptoms, and rater bias.

The drug-placebo difference in a trial with compromised blinding therefore reflects:

  1. Drug-specific pharmacological effects (the intended target)
  2. Differential expectancy effects (patients who recognize side effects know they are on active drug and may experience enhanced placebo effects)
  3. Rater bias (raters who recognize side effects may rate those patients as more improved)

The magnitude of components (2) and (3) depends on how easily the drug can be distinguished from placebo. For drugs with subtle or rare side effects, unblinding may be minimal. For drugs with frequent, recognizable side effects—appetite suppression, sexual dysfunction, dry mouth, sleep disturbance—unblinding is substantial.

1.2 Recognizable side effects in psychiatric drug trials

Stimulants (methylphenidate, amphetamines):

The Cochrane systematic review of methylphenidate for ADHD (Storebø et al., 2015) found that among trials comparing methylphenidate to placebo, methylphenidate produced:

  • Decreased appetite: RR 3.66 (95% CI 2.56–5.23); 16 trials, 2,962 participants. This is a 266% increase in risk.
  • Sleep problems: RR 1.60 (95% CI 1.15–2.23); 13 trials, 2,416 participants. This is a 60% increase in risk.

These side effects are frequent (occurring in 15–30% of children on methylphenidate) and easily recognized by parents, teachers, and clinicians. The Cochrane authors note: "Due to the frequency of non-serious adverse events associated with methylphenidate, the blinding of participants and outcome assessors is particularly challenging. To accommodate this challenge, an active placebo should be sought and utilised."

In one laboratory-classroom study (Wigal et al., 2017), the most common adverse events for methylphenidate extended-release were decreased appetite (reported in multiple children), upper abdominal pain, headache, and insomnia. These are not subtle effects.

Antidepressants (SSRIs and tricyclics):

SSRIs produce recognizable side effects including:

  • Sexual dysfunction (delayed orgasm, reduced libido) in 40–65% of patients [Montejo et al., 1997, 2001; Rothschild, 2000]
  • Gastrointestinal disturbance (nausea, diarrhea) especially in the first 2 weeks
  • Activation or sedation depending on the specific SSRI

Tricyclic antidepressants produce anticholinergic effects:

  • Dry mouth, constipation, urinary retention
  • Sedation, weight gain
  • Orthostatic hypotension

Rief et al. found that the greater tolerance of SSRIs versus tricyclics was identical in both placebo and active groups, indicating that awareness of expected side effects influences tolerability independent of actual drug exposure.

1.3 Evidence that side effects compromise blinding

Direct evidence from guessing studies:

In a double-blind trial of phenylpropanolamine (PPA) for obesity, Moscucci et al. (1987) asked participants at study end to guess their treatment assignment. The results:

  • 74% of placebo participants correctly guessed placebo
  • 43% of PPA participants correctly guessed PPA (57% guessed incorrectly or didn't know)

The most frequently reported basis for guessing PPA was appetite control, even among placebo participants. The most frequently reported basis for guessing placebo was lack of adverse drug reactions, even among PPA participants.

Critically, participants receiving either PPA or placebo who guessed PPA lost more weight, had less diet difficulty, and had more adverse drug reactions than participants receiving either PPA or placebo who guessed placebo. Outcome correlated with perceived assignment, not actual assignment.

Indirect evidence from active-placebo trials:

Two antidepressant trials tested the integrity of the blind by asking raters to guess medication group (Uhlenhuth et al., 1964; Weintraub et al., 1963). Guesses were more accurate than chance, though not statistically significant in either trial. In the Weintraub trial, both raters assessed those they guessed to be on active drug as more improved, regardless of actual assignment.

One trial (Hollister et al., 1964) reported that side effects had been more prominent in patients on antidepressants, indicating that residual unblinding effects may have occurred despite the use of active placebos containing atropine.

1.4 The mechanism: expectancy enhancement

When a patient experiences a recognizable side effect and infers (correctly or incorrectly) that they are receiving active medication, their expectancy of improvement increases. This enhanced expectancy can then produce symptom improvement independent of the drug's pharmacological mechanism.

Rutherford et al. (2017) demonstrated this experimentally. In a randomized trial, adult outpatients with major depressive disorder were assigned to either:

  1. Open citalopram (100% probability of receiving active drug)
  2. Placebo-controlled citalopram (50% probability of receiving active drug)

Postrandomization expectancy scores were significantly higher in the open group (mean 12.1, SD 2.1) versus the placebo-controlled group (mean 11.0, SD 2.0). Hamilton Depression Rating Scale scores declined faster in the open group. Patient expectations postrandomization partially mediated group effects on week-8 HRSD scores.

This finding has a direct implication for double-blind trials: if merely knowing one's probability of receiving active treatment (100% versus 50%) affects outcomes, then correctly guessing one's assignment mid-trial due to side effects would produce a similar expectancy-mediated enhancement.

Translation: In a standard pill-placebo trial, patients who experience recognizable side effects and infer they are on active drug effectively transition from the 50%-probability condition to a high-confidence active-drug condition. Their outcomes then reflect both the drug's pharmacology and the enhanced expectancy produced by perceived assignment.


Part II — Active placebos: what happens when side effects are mimicked

2.1 The Moncrieff meta-analysis of active-placebo trials

To address the unblinding problem, some early antidepressant trials used active placebos containing atropine, which produces anticholinergic effects (dry mouth, constipation) similar to those of tricyclic antidepressants. The rationale was that if patients and raters in both groups experience similar side effects, the blind would be better maintained.

Moncrieff et al. (2004) conducted a Cochrane systematic review of antidepressants versus active placebos. Nine trials involving 751 participants were included. All compared tricyclic antidepressants with active placebos containing atropine. A minimum dose of 100 mg amitriptyline or equivalent was used in all but one trial.

Individual trial results:

  • Daneman (1961): SMD 1.1 (95% CI 0.8–1.4) in favor of imipramine. This trial showed a large positive effect, but closer inspection revealed possible problems with blinding and selective reporting. The placebo response rate was unusually poor (9% at 8 weeks).

  • Hussain (1970): SMD 0.79 (95% CI 0.09–1.5). The trial indicated antidepressants were superior, though the authors found no significant difference using categorical analysis.

  • Weintraub (1963): Results from two raters were inconsistent. Hospital director: SMD 0.14 (95% CI –0.34 to 0.62). Ward doctor: SMD 0.63 (95% CI 0.15–1.11).

  • Uhlenhuth (1964): Unadjusted SMD 0.60 (95% CI 0.02–1.2). However, when adjusted for substantial baseline differences in depression severity, SMD 0.35 (95% CI –0.25 to 0.96)—no longer significant.

  • Hollister (1964): SMD 0.19 (95% CI –0.24 to 0.63). No difference between tricyclics and active placebo.

  • Friedman (1966): SMD 0.13 (95% CI –0.37 to 0.64). No difference.

  • Friedman (1975): SMD 0.14 (95% CI –0.14 to 0.42). No difference.

  • Murphy (1984): SMD –0.36 (95% CI –1.0 to 0.28). No difference (all participants also received cognitive therapy).

  • Wilson (1963): SMD –0.26 (95% CI –1.10 to 0.58). No difference.

Pooled analysis:

All nine trials combined: SMD 0.39 (95% CI 0.24–0.54). This is a highly significant difference, but there was high heterogeneity (χ² = 36.3, df = 8, p < 0.001). The source of heterogeneity was the Daneman (1961) trial.

Excluding Daneman (1961): SMD 0.17 (95% CI 0.00–0.34). Heterogeneity reduced to non-significant levels (χ² = 8.51, df = 7, p = 0.29).

Subgroup analysis by setting:

  • Inpatient trials (predominantly severe/endogenous depression): SMD 0.12 (95% CI –0.14 to 0.38). Not statistically significant.
  • Outpatient trials (predominantly moderate/neurotic depression): SMD 0.52 (95% CI 0.34–0.70), but excluding Daneman reduced this to SMD 0.20 (95% CI –0.02 to 0.43).

Interpretation:

The more conservative estimates—excluding the problematic Daneman trial—suggest that the difference between antidepressants and active placebos is small and possibly not statistically significant (SMD 0.17, 95% CI 0.00–0.34). For inpatient trials, the effect is clearly non-significant (SMD 0.12).

Comparison with inert-placebo trials: Meta-analyses of antidepressants versus inert placebos from the same era typically found effect sizes in the range SMD 0.4–0.8. The Moncrieff analysis, using active placebos, found SMD 0.17–0.39 depending on which trials were included.

Conclusion: Moncrieff et al. write: "The more conservative estimates from the present analysis found that differences between antidepressants and active placebos were small. This suggests that unblinding effects may inflate the efficacy of antidepressants in trials using inert placebos."

2.2 Quality and limitations of the active-placebo evidence

The Moncrieff analysis has important limitations:

  1. Age of trials: All trials were conducted in the 1960s–1980s before standardized diagnostic criteria (DSM-III) and outcome measures (Hamilton Rating Scale) were widely adopted. Outcome measures were a mixture of validated instruments and categorical ratings.

  2. Small sample sizes: Individual trials ranged from 20–200 participants. Power to detect effects was limited.

  3. Heterogeneity in methods: Some trials reported only categorical improvement ratings, which had to be converted to continuous data. Standard deviations were estimated from other trials in some cases.

  4. Limited generalizability: All trials used tricyclic antidepressants. No active-placebo trials of SSRIs or SNRIs were included, because these drugs were not widely used when active-placebo designs were feasible.

  5. Residual unblinding: Even with active placebos, one trial (Hollister 1964) noted that side effects were more prominent in the antidepressant group, suggesting incomplete blinding.

Despite these limitations, the pattern is consistent: when side effects are mimicked in the placebo group, the measured drug-placebo difference shrinks substantially.

2.3 Why active placebos are no longer used

Active placebos present ethical and practical challenges:

Ethical: Administering a substance that produces side effects but no therapeutic benefit violates the principle of nonmaleficence. Modern institutional review boards are unlikely to approve such designs.

Practical: Selecting an active placebo that mimics the side-effect profile of the study drug without producing therapeutic effects is difficult. Atropine mimics anticholinergic effects of tricyclics, but what would mimic SSRI sexual dysfunction without affecting serotonin? There is no obvious candidate.

Blinding imperfect anyway: Even with active placebos, differences in the frequency or intensity of side effects can reveal treatment assignment. If 30% of antidepressant recipients experience dry mouth and 15% of atropine-placebo recipients do, the blind is partially compromised.

For these reasons, active-placebo designs have largely disappeared from the psychiatric literature. Modern trials use inert placebos and hope that blinding holds—despite ample evidence that it often does not.


Part III — The severity-by-placebo interaction: placebo falls, drug stays flat

3.1 The Kirsch 2008 FDA analysis

Kirsch et al. (2008) obtained FDA data for four new-generation antidepressants: fluoxetine, venlafaxine, nefazodone, and paroxetine. The dataset comprised 35 clinical trials involving 5,133 patients (3,292 randomized to medication, 1,841 to placebo). Trial duration ranged from 4–8 weeks, with most lasting 6 weeks.

Overall drug-placebo difference:

Weighted mean improvement was 9.60 points on the Hamilton Rating Scale (HRSD) in drug groups and 7.80 points in placebo groups, yielding a mean difference of 1.80 points. Represented as standardized mean difference (Cohen's d), mean change was d = 1.24 for drug and d = 0.92 for placebo, a difference of d = 0.32.

The NICE criterion for clinical significance is a 3-point HRSD difference or d = 0.50. The overall drug-placebo difference (1.80 points, d = 0.32) falls below this threshold.

Baseline severity as a moderator:

The relationship between baseline severity and improvement was assessed using meta-regression. The key findings:

  1. Drug groups showed a curvilinear (∩-shaped) relationship: Improvement increased with severity up to a point, then declined slightly at the highest severity levels.

  2. Placebo groups showed a strong negative linear relationship: Placebo response fell sharply as baseline severity increased.

  3. The drug-placebo difference increased with severity, but the mechanism was decreased placebo response, not increased drug response.

The drug-placebo difference reached the NICE criterion for clinical significance (d ≥ 0.50 or ≥ 3 HRSD points) only for patients with baseline HRSD scores above 28—the upper end of the very severely depressed category.

Visual summary (from Figure 2 of Kirsch et al., 2008):

At baseline HRSD 18 (moderate depression): drug improvement ≈ 10 points, placebo improvement ≈ 10 points. Difference ≈ 0 points.

At baseline HRSD 23 (severe depression): drug improvement ≈ 11 points, placebo improvement ≈ 8 points. Difference ≈ 3 points (borderline clinical significance).

At baseline HRSD 28 (very severe depression): drug improvement ≈ 11 points, placebo improvement ≈ 6 points. Difference ≈ 5 points (clinically significant).

At baseline HRSD 32 (very severe depression, upper range): drug improvement ≈ 10 points, placebo improvement ≈ 4 points. Difference ≈ 6 points.

The critical inference: The widening gap is not because drugs work better in severe depression. Drug response is relatively flat across severity levels (10–11 points improvement). The gap widens because placebo response drops from 10 points at moderate severity to 4 points at very high severity.

3.2 Mechanisms and implications

Why does placebo response drop in severe depression?

Kirsch et al. speculate that severely depressed patients may have difficulty generating or maintaining positive expectations, which are thought to mediate placebo effects. Cognitive symptoms (hopelessness, anhedonia, concentration difficulties) may interfere with expectancy formation.

Alternatively, severely depressed patients may have less room for "response to context" because vegetative symptoms (sleep disturbance, psychomotor retardation) dominate the clinical picture and are less responsive to psychological mechanisms.

What this means for treatment decisions:

  1. For patients with moderate depression (HRSD < 23): The drug-placebo difference is clinically insignificant. Antidepressants produce improvement, but so does placebo, and the difference is less than 3 HRSD points. Other factors (cost, side effects, patient preference, availability of psychotherapy) may weigh more heavily in treatment selection.

  2. For patients with very severe depression (HRSD > 28): The drug-placebo difference is clinically significant. Antidepressants provide meaningful benefit over placebo, driven by the collapse of placebo response rather than enhanced drug response.

  3. For patients in between: The drug-placebo difference is in a grey zone. Individual variation likely exceeds the mean difference.

3.3 Critiques and rebuttals

[contested] This analysis has been criticized on the grounds that:

  1. The four drugs studied (especially nefazodone, which was later withdrawn) are not representative of the full antidepressant class.

  2. The Cipriani et al. (2018) network meta-analysis of 522 trials found that all 21 antidepressants were more effective than placebo (odds ratios 1.37–2.13), with no evidence of a severity-by-treatment interaction.

Rebuttal: The Cipriani analysis pooled across all severity levels without testing for moderation by baseline severity. If trials with mixed severity levels are combined, a severity interaction can be masked. The two analyses address different questions: Kirsch tests whether the drug-placebo difference varies by severity (it does), while Cipriani tests whether all drugs beat placebo on average (they do). Both can be true.

A more substantive critique is that the Kirsch dataset is limited to four drugs and 35 trials. Replication with a larger, more diverse dataset would strengthen the finding. To date, such replication has not been published.

What survives: The pattern Kirsch documents—widening drug-placebo difference driven by falling placebo response—is internally consistent across the 35 trials in the FDA dataset. Whether it generalizes to all antidepressants or all severity ranges remains an open question, but the finding is not an artifact of selective reporting (because the dataset includes all FDA-registered trials, published and unpublished).


Part IV — Publication bias as a separate inflation pathway

4.1 The Turner 2008 analysis

Turner et al. (2008) obtained FDA reviews for studies of 12 antidepressant agents approved between 1987 and 2004, involving 12,564 patients. They identified 74 phase 2 and 3 randomized, double-blind, placebo-controlled trials. They then conducted a systematic literature search to identify published reports of those trials.

Key findings:

OutcomeFDA dataPublished literature
Number of studies7451 published
Studies deemed positive by FDA38 (51%)37 published
Studies deemed negative/questionable by FDA36 (49%)3 published as negative
11 published as positive
22 not published
Apparent success rate51%94%

Studies viewed by the FDA as positive were approximately 12 times as likely to be published in agreement with the FDA analysis as studies with non-positive results (risk ratio 11.7, 95% CI 6.2–22.0, p < 0.001).

Effect-size inflation:

For each of the 12 drugs, the effect size derived from published literature was higher than the effect size derived from the full FDA dataset. The increase in effect size ranged from 11–69% for individual drugs and was 32% overall.

The mean weighted effect size was:

  • FDA data (all 74 trials): g = 0.31 (95% CI 0.27–0.35)
  • Published literature (51 trials): g = 0.41 (95% CI 0.36–0.45)

The 32% inflation (from 0.31 to 0.41) is the difference between the true effect (FDA data) and the apparent effect (published literature).

Breakdown by publication status within FDA data:

  • Published studies: mean effect size g = 0.37 (95% CI 0.33–0.41)
  • Unpublished studies: mean effect size g = 0.15 (95% CI 0.08–0.22)

The difference between published and unpublished studies was significant (p < 0.001).

4.2 Mechanisms of publication bias

Turner et al. conclude: "We cannot determine whether the bias observed resulted from a failure to submit manuscripts on the part of authors and sponsors, from decisions by journal editors and reviewers not to publish, or both."

Both mechanisms are plausible:

Sponsor non-submission: Pharmaceutical companies may choose not to submit negative trials for publication, either because journal reviewers are unlikely to accept them or because publication of negative results could harm the drug's commercial prospects.

Editorial rejection: Journal editors may preferentially accept trials with positive results because they are perceived as more interesting or because negative results are harder to interpret (failure to find an effect could reflect insufficient power, poor trial conduct, or true lack of efficacy).

Spin in published reports: Eleven trials deemed negative or questionable by the FDA were published in a way that conveyed a positive outcome. This involved subordinating non-significant results on the prespecified primary outcome and highlighting positive results on secondary outcomes, or omitting the non-significant primary outcome entirely.

4.3 Implications for evidence-based medicine

The Turner analysis demonstrates that the published literature on antidepressants provides a systematically biased view of drug efficacy. A clinician reading only published trials would conclude that 94% of trials are positive, when the FDA data show that 51% are positive. Effect sizes are inflated by nearly one-third.

Is this a historical problem? The Turner dataset covers trials from 1987–2004. Since then, trial registration requirements have improved transparency. ClinicalTrials.gov was established in 2000, and the FDA Amendments Act (2007) mandated registration of all phase 2–4 trials.

However, the Cipriani et al. (2018) network meta-analysis—covering 522 trials up to 2016—still found that 9% of trials were rated as high risk of bias and 73% as moderate risk, with certainty of evidence ranging from moderate to very low. Publication bias has not been eliminated, though it may have been reduced.

The structural lesson: Publication bias is a predictable consequence of (1) financial incentives to publish positive results, (2) editorial preferences for novel or positive findings, and (3) lack of consequences for non-publication. Correcting it requires mandatory public registration of all trials and penalties for non-publication, not voluntary compliance.


Part V — Expectancy and nocebo: the other side of the coin

5.1 Nocebo effects in placebo arms

Nocebo effects—adverse events occurring after administration of an inert substance—are substantial in psychiatric trials. Participants in placebo arms report side effects at rates closely resembling those in active-drug arms, including drug-specific effects.

Antidepressant examples:

In placebo-controlled SSRI trials, placebo-arm participants report:

  • Sexual dysfunction (though at lower rates than the active-drug arm)
  • Nausea and gastrointestinal symptoms
  • Activation or sedation

Rief et al. (2009) conducted a systematic review and meta-analysis of 143 placebo-controlled trials (12,742 patients). They found that the greater tolerance of SSRIs versus tricyclics was identical in both placebo and active groups. Far more adverse effects were reported in TCA-placebo groups than in SSRI-placebo groups: dry mouth (OR 3.5, 95% CI 2.9–4.2), drowsiness (OR 2.7, 95% CI 2.2–3.4), constipation (OR 2.7, 95% CI 2.1–3.6), and sexual problems (OR 2.3, 95% CI 1.5–3.5). This indicates that awareness of expected side effects influences tolerability independent of drug exposure.

Conditioning of nocebo effects:

Rheker et al. (2017) conducted a randomized, controlled, double-blind study in which healthy volunteers received 50 mg amitriptyline for 4 nights together with a novel-tasting drink as a conditioned stimulus. After a 3-day washout, participants received placebo with the same drink. Placebo administration reproduced amitriptyline-specific side effects, including dry mouth, dizziness, and circulation-related problems. This demonstrates that nocebo responses can be conditioned through prior drug exposure via classical conditioning.

Prevalence: Estimates of nocebo-related adverse effects in antidepressant trials range from 44.7% (Mitsikostas et al.) to 63.7% (Dodd et al.). These figures apply to participants receiving placebo who report at least one adverse event.

5.2 Expectancy as a mediator of treatment outcomes

The Rutherford et al. (2017) study, discussed earlier, demonstrated that patient expectancy mediates outcomes. Key findings:

  • Participants randomized to open citalopram (100% probability of active drug) had higher postrandomization expectancy scores than participants randomized to placebo-controlled citalopram (50% probability).
  • HAM-D scores declined faster in the open group.
  • Postrandomization expectancy partially mediated the group effect on week-8 HAM-D scores.

This shows that merely knowing one's probability of receiving active treatment affects symptom trajectories, independent of actual drug assignment.

Implications for double-blind trials: In a standard double-blind trial, all participants start with 50% probability of active drug. But if a participant experiences a recognizable side effect (e.g., sexual dysfunction on an SSRI), they update their subjective probability upward—perhaps to 80% or 90%. This enhanced expectancy then produces symptom improvement via psychological mechanisms (conditioning, attention, interpretation of symptoms, motivation).

The drug-placebo difference in such a trial reflects:

  1. Pharmacological effects of the drug
  2. Differential expectancy effects (higher in the group that correctly infers active treatment)

Disentangling these requires either active placebos (to equalize side effects and thus expectancy) or measures of expectancy throughout the trial (to statistically control for it).

5.3 Expectancy effects vary by disorder and baseline characteristics

Depression and nocebo vulnerability: Cognitive biases characteristic of depression—sustained attention to negative information, interpretation of ambiguous information as negative, preferential recall of negative memories—may amplify nocebo effects. Depressed patients are primed to expect and notice negative outcomes, including side effects.

However, when comparing depression to other brain diseases (motor neuron disease, Parkinson's, Alzheimer's), all of the latter show more prevalent nocebo adverse events than depression. This suggests that nocebo vulnerability is not unique to depression but may be elevated in conditions with prominent cognitive symptoms.

Role of information provided: Informed consent documents list potential side effects of the study drug. Participants who read that "30% of patients experience sexual dysfunction" may then monitor their sexual functioning more closely and interpret normal fluctuations as drug-induced side effects. This is a form of nocebo effect driven by attention and attribution.

One strategy to mitigate this is to provide balanced information about placebo nocebo effects—explaining that people in placebo arms also experience side effects, and that side effects do not necessarily indicate active drug assignment. However, ethical requirements for informed consent limit how much information can be withheld.


Part VI — What the evidence does and does not establish

6.1 Established

  1. Recognizable side effects compromise blinding. In ADHD stimulant trials, decreased appetite and sleep problems occur 2–4 times more often on active drug than placebo. In antidepressant trials, SSRI-specific side effects are easily recognized. When patients and raters guess treatment assignment based on side effects, outcome ratings correlate with perceived assignment independent of actual assignment.

  2. Active placebos reduce the measured drug-placebo difference. Moncrieff et al. (2004) found SMD 0.17 (95% CI 0.00–0.34) for antidepressants versus active placebos, compared to SMD 0.4–0.8 in trials using inert placebos from the same era. For inpatient trials (severe depression), the effect was SMD 0.12 (95% CI –0.14 to 0.38)—not statistically significant. This suggests that unblinding inflates efficacy estimates in inert-placebo trials.

  3. The drug-placebo difference is severity-dependent. Kirsch et al. (2008) found that the difference reaches clinical significance only at HRSD > 28. The mechanism is reduced placebo response in severe depression (from 10 points at moderate severity to 4 points at very high severity), not increased drug response (which remains 10–11 points across severity levels).

  4. Publication bias inflated efficacy by 32%. Turner et al. (2008) documented that 31% of FDA-registered trials were not published, and 11 additional trials were published with positive spin despite negative FDA conclusions. Effect sizes in the published literature were 32% higher than in the full FDA dataset.

  5. Patient expectancy mediates placebo effects. Rutherford et al. (2017) showed that knowing one's probability of receiving active treatment (100% versus 50%) affects symptom trajectories via expectancy.

  6. Nocebo effects are substantial. Placebo-arm participants report side effects at rates resembling active-drug arms. Nocebo effects can be conditioned through prior drug exposure.

  7. Control condition type affects remission rates. Casacalenda et al. (2002) found remission rates of 46.4% for medication, 46.3% for psychotherapy, and 24.4% for control conditions (pill placebo plus clinical management, treatment as usual by general physician, and low-intensity supportive therapy) in six trials of outpatients with mild-to-moderate depression.

6.2 Not established

That psychiatric drug effects are entirely placebo: Even after correcting for publication bias and using active placebos, a residual drug-placebo difference remains. It is smaller than the published literature suggests, and it may not be clinically significant at moderate severity levels, but it is not zero.

That double-blind designs cleanly isolate drug-specific effects: Blinding is compromised by recognizable side effects. The drug-placebo difference includes pharmacological effects, differential expectancy, and rater bias. Disentangling these requires active placebos or expectancy measures, neither of which are standard.

That placebo effects are inert or trivial: In both drug and placebo arms, most improvement is attributable to placebo effects (defined broadly as response to context, expectancy, natural history, regression to the mean). Drug-specific effects are additive on top of large placebo effects.

That active placebos are a complete solution: Active placebos reduce but do not eliminate expectancy effects. Ethical concerns and difficulty in selecting appropriate active substances prevent widespread use.

Which patients will experience clinically meaningful drug-specific effects: Baseline severity predicts the magnitude of the drug-placebo difference on average, but individual variation is large. At moderate severity, some patients respond better to drug than placebo, and others respond equally or better to placebo. There is no validated algorithm for predicting individual response.

6.3 The two overclaims, and what replaces them

Overclaim 1: "It's all placebo."

This is false. After correcting for publication bias, using active placebos, and restricting analysis to severe depression, a drug-placebo difference remains. It is smaller and less consistent than marketing suggests, but it is not zero.

Overclaim 2: "Double-blind trials cleanly isolate drug-specific effects."

This is also false. Recognizable side effects compromise blinding. The drug-placebo difference reflects pharmacology plus differential expectancy plus rater bias. The relative contribution of each component varies by drug, trial design, and baseline severity, but in trials with frequent recognizable side effects, expectancy contamination is substantial.

What replaces both:

Pill-placebo trials compare active-drug-with-expectancy against inert-substance-with-uncertain-expectancy in a context where side effects often reveal treatment assignment. The drug-placebo difference captures:

  1. Drug-specific pharmacological effects (target of interest)
  2. Differential expectancy effects due to unblinding (confound)
  3. Publication bias inflation (when reading the literature)

Disentangling (1) from (2) requires better blinding strategies—active placebos, expectancy measures, or novel trial designs that reduce the signal-to-noise ratio of side effects. Correcting for (3) requires complete trial registration and penalties for non-publication.

Neither of these is standard practice. Until they are, the drug-placebo differences reported in the literature should be understood as upper bounds, not clean estimates.


Conclusion

Placebo-controlled trials are the foundation of evidence-based medicine, but they are not inert-versus-active comparisons. They are context-with-pharmacology versus context-with-uncertainty comparisons, and when recognizable side effects compromise blinding, the uncertainty collapses.

The drug-placebo difference in such trials includes genuine pharmacological effects, but it also includes expectancy amplification, rater bias, and—when reading the published literature—a 32% inflation from selective publication.

Active-placebo trials reduce this contamination by mimicking side effects, but they show much smaller drug-placebo differences (SMD 0.17 versus 0.39). For severe depression, the difference may not be statistically significant at all.

The Kirsch severity-by-placebo interaction is real: placebo response drops from 10 HRSD points at moderate severity to 4 points at very high severity, while drug response remains flat at 10–11 points. The implication is that at moderate severity, the drug-placebo difference is clinically insignificant—not because drugs don't work, but because placebo works nearly as well.

None of this supports the radical claim that psychiatric drugs are inert. Corrected estimates remain positive. But it does refute the orthodox claim that double-blind designs cleanly isolate drug-specific effects. They do not. Expectancy and unblinding contaminate the estimate, and publication bias inflates it further.

The honest position: psychiatric drugs produce modest benefits over placebo, driven partly by pharmacology and partly by expectancy. At moderate severity, the benefits may not be clinically meaningful. At very high severity, they are. The mechanisms are incompletely understood, and individual response is highly variable. The published literature overstates efficacy by approximately one-third due to selective reporting.

Neither "it's all placebo" nor "double-blind means the estimate is clean" survives contact with the data.

About the author

Paul Stephen

Founder, Apatheia Labs

Evidence-governed research publication — Prosoche applied in the open.

All essays

Follow

New work, in your reader

New essays and audits publish to RSS — no inbox, no list. Point your reader at the feed and they arrive as they land.

Subscribe via RSS

Published by Apatheia Labs. All rights reserved. Quote freely with attribution; redistribute with permission.