Evidence Review
Psychotherapy Comparators and the Evidence
Why the control condition determines what you find
On this page25 sections
Companion reviews: Antipsychotics and the Evidence, Placebo, expectancy, and unblinding, Depression — antidepressants and the evidence, ADHD — stimulants and the evidence, Anxiety — a critical review of the evidence, and Diagnostic Thresholds, Prevalence, and Overdiagnosis.
Executive summary
Two folk models dominate public discussion of psychotherapy efficacy: that psychotherapy produces large effects for most mental disorders (the optimistic position), and its inversion, that psychotherapy is merely "talking to a friend" and no better than any other supportive intervention (the dismissive position). Both are overclaims, and the evidence supports neither.
Six things the evidence actually establishes:
-
Waitlist control groups inflate psychotherapy effect sizes compared to care-as-usual or pill placebo controls. In depression trials, psychotherapy versus waitlist produces effect sizes of Hedges' g = 0.95 (95% CI: 0.85–1.04), while psychotherapy versus care-as-usual produces g = 0.63 (95% CI: 0.55–0.71; Cuijpers et al., 2024). The difference is driven by reduced improvement in waitlist conditions (g = 0.37 pre-post, 95% CI: 0.28–0.46) compared to care-as-usual conditions (g = 0.64 pre-post, 95% CI: 0.50–0.78). Across anxiety disorders, PTSD, and OCD, 71–96% of trials use waitlist controls (GAD 70.8% to specific phobia 95.8%); the overall waitlist share across all eight disorders is 61%. For depression, only 42% use waitlist.
-
Absolute response rates for psychotherapy are modest across disorders, ranging from 24% to 42%. In a meta-analysis of 441 trials (33,881 patients) comparing psychotherapy against control conditions (waitlist, care-as-usual, or pill placebo), response rates (≥50% symptom reduction) were: major depressive disorder 0.42 (95% CI: 0.39–0.45); PTSD 0.38 (95% CI: 0.33–0.43); panic disorder 0.38 (95% CI: 0.33–0.43); GAD 0.36 (95% CI: 0.30–0.42); social anxiety disorder 0.32 (95% CI: 0.29–0.37); OCD 0.38 (95% CI: 0.30–0.47); borderline personality disorder 0.24 (95% CI: 0.15–0.36) (Cuijpers et al., 2024, World Psychiatry). Control-condition response rates ranged from 0.05 (OCD) to 0.19 (depression). Most patients receiving psychotherapy do not achieve 50% symptom reduction.
-
In direct head-to-head trials, psychotherapy and medication produce nearly identical remission rates in outpatients with mild-to-moderate depression. Casacalenda et al. (2002) identified six randomized trials (883 patients) comparing tricyclic antidepressants or phenelzine, psychotherapy (primarily CBT and IPT), and control conditions. Intent-to-treat remission rates were 46.4% for medication, 46.3% for psychotherapy, and 24.4% for control conditions (p < 0.0001 for both active treatments versus controls; no significant difference between active treatments). Control conditions were pill placebo plus clinical management (3 trials), treatment-as-usual by a general physician (2 trials, with ~45% receiving antidepressants), and low-intensity supportive therapy (1 trial).
-
Type of control condition significantly affects measured efficacy, even within the same disorder and psychotherapy type. For depression, response rates differed significantly by control type (p values not reported but subgroup analyses significant for MDD, panic disorder, and PTSD; Cuijpers et al., 2024). The mechanism is differential pre-post change in the control arm, not differential change in the therapy arm. Waitlist participants improve less (g = 0.37) than care-as-usual participants (g = 0.64), while therapy-arm pre-post effects do not differ significantly by control type.
-
Therapist effects account for a larger share of outcome variance than treatment type effects. In a meta-analysis of alliance-outcome studies (nearly 200 studies, >14,000 patients), the aggregate alliance-outcome correlation was r = 0.27 (equivalent to d = 0.57; Wampold, 2015). Multilevel analyses show that therapist contributions to alliance predict outcome, while patient contributions do not: more effective therapists form strong alliances across a range of patients, including those with poor attachment histories. Therapist empathy, positive regard, and genuineness correlate with outcome at d = 0.63, 0.56, and 0.49 respectively.
-
The dodo-bird claim of complete psychotherapy equivalence is contested and depends heavily on which comparisons are included. Wampold and colleagues argue that bona fide psychotherapies (those delivered with a cogent rationale, agreed-upon tasks, and therapist belief in the treatment) produce equivalent outcomes because they activate common factors (alliance, empathy, expectations, therapist differences) rather than specific-ingredient effects. Critics note that many meta-analyses showing equivalence pool trials with inadequate statistical power to detect treatment differences, include therapies without established efficacy for the disorder in question, and fail to account for researcher allegiance. Subgroup analyses in Cuijpers et al. (2024) found significant response-rate differences by therapy type for depression (p < 0.001), panic disorder (p = 0.02), specific phobia (p = 0.003), and PTSD (p < 0.001), but not for social anxiety, GAD, OCD, or BPD.
What the evidence does not establish:
- That psychotherapy produces large effects for most people (it does not; response rates are 24–42%)
- That psychotherapy is equivalent to "talking to a friend" (informal support lacks the structure, alliance-building, cogent explanation, and health-promoting actions that characterize psychotherapy)
- That all bona fide psychotherapies are equivalent (treatment differences exist for some disorders but not others)
- That specific ingredients are irrelevant (the evidence is mixed; some disorders show treatment differences, others do not)
- Which patients will achieve clinically meaningful improvement versus non-response
The structural problem that conditions all of this: Psychotherapy efficacy depends on what you compare it to. Waitlist controls inflate effect sizes by restricting spontaneous recovery; care-as-usual controls produce smaller effects because participants receive some active support; pill-placebo controls are rarely used (5% of trials) and confound expectancy effects with drug-specific pharmacology. Any efficacy claim that does not specify (1) the control condition, (2) whether the control was waitlist, care-as-usual, pill placebo, or psychological placebo, and (3) whether response, remission, or effect size is reported, should be mentally adjusted downward.
How to read this review
The single most important heuristic
An effect size or response rate without a named comparator is not a clean estimate. Psychotherapy versus waitlist produces g = 0.95 for depression; psychotherapy versus care-as-usual produces g = 0.63. Psychotherapy versus pill placebo is reported in only 5% of trials across disorders. Any claim that "psychotherapy works" must specify compared to what.
Three structural problems
Control-condition heterogeneity inflates aggregate effect sizes. Meta-analyses pooling waitlist, care-as-usual, and pill-placebo controls report larger effects than meta-analyses restricted to care-as-usual or pill-placebo controls. Waitlist participants improve less (g = 0.37 pre-post) than care-as-usual participants (g = 0.64 pre-post), while therapy-arm improvement does not differ by control type. The inflation is not due to psychotherapy working better against waitlist; it is due to waitlist producing less spontaneous recovery.
Small absolute response rates are obscured by large relative risks. A response rate of 0.42 for psychotherapy versus 0.19 for controls produces a relative risk of 2.09—"more than twice as likely to respond." But 58% of therapy recipients still do not achieve 50% symptom reduction, and 81% of control participants do not respond. Relative risks magnify modest absolute effects. Number needed to treat (NNT) is more transparent: for depression, NNT = 4.8, meaning five patients must be treated for one additional response beyond the control rate.
The dodo-bird debate conflates two distinct questions. (1) Are all bona fide psychotherapies equivalent on average across disorders? (Wampold: yes, because common factors dominate; critics: no, because many included therapies lack established efficacy for specific disorders.) (2) Do specific therapeutic ingredients produce differential effects within a disorder? (The evidence is mixed: treatment differences exist for depression, panic, specific phobia, and PTSD; no differences for social anxiety, GAD, OCD, or BPD.) The first question is a claim about general mechanisms; the second is a claim about disorder-specific treatment selection.
Conventions
Claims are marked [contested] where the literature genuinely disagrees and [unverified] where a figure is widely repeated but could not be traced to a primary source in preparing this review.
Confidence intervals are omitted where they could not be verified against the primary text, rather than reconstructed.
Part I — Comparator type determines measured efficacy
1.1 Waitlist versus care-as-usual: the 50% inflation
Cuijpers et al. (2024, Epidemiology and Psychiatric Sciences) conducted a meta-analysis comparing psychotherapy trials using waitlist controls against trials using care-as-usual controls for adult depression. They included 333 randomized trials (472 comparisons; 41,480 participants): 141 trials with waitlist controls, 195 with care-as-usual controls, and 3 with both.
Main findings:
The overall effect size (therapy versus control at post-test) was g = 0.77 (95% CI: 0.71–0.84) with high heterogeneity (I² = 84%). Subgroup analysis revealed:
- Waitlist-controlled trials: g = 0.95 (95% CI: 0.85–1.04; I² = 80%)
- Care-as-usual-controlled trials: g = 0.63 (95% CI: 0.55–0.71; I² = 85%)
- Difference: p < 0.001
This difference remained significant in all sensitivity analyses, including a meta-regression adjusting for all measured differences between waitlist and care-as-usual trials (therapy type, treatment format, recency, target group, recruitment strategy, number of treatment arms, number of outcome measures).
The mechanism is reduced spontaneous recovery in waitlist, not enhanced therapy response:
Pre-post effect sizes within control conditions:
- Waitlist: g = 0.37 (95% CI: 0.28–0.46)
- Care-as-usual: g = 0.64 (95% CI: 0.50–0.78)
- Difference: significant
Pre-post effect sizes within therapy conditions showed no significant difference between waitlist-controlled and care-as-usual-controlled trials.
Translation: Waitlist inflates measured psychotherapy efficacy by suppressing control-arm improvement, not by enhancing therapy-arm improvement. Participants assigned to waitlist improve less than participants assigned to care-as-usual, even though both groups are untreated by the experimental psychotherapy. The 50% larger effect size in waitlist-controlled trials (g = 0.95 versus g = 0.63) reflects this control-arm difference.
1.2 The prevalence of waitlist controls varies by disorder
Cuijpers et al. (2024, World Psychiatry) examined 441 trials (33,881 patients) across eight mental disorders. Control condition types by disorder:
| Disorder | Waitlist (%) | Care-as-usual (%) | Pill placebo (%) |
|---|---|---|---|
| Depression | 42.0 | 54.9 | 3.1 |
| Panic disorder | 75.0 | 12.5 | 12.5 |
| Social anxiety | 84.6 | 7.7 | 7.7 |
| GAD | 70.8 | 22.9 | 6.2 |
| Specific phobia | 95.8 | 4.2 | 0 |
| PTSD | 79.3 | 18.8 | 3.9 |
| OCD | 72.7 | 18.2 | 9.1 |
| BPD | 0 | 100.0 | 0 |
Key points:
-
Anxiety disorders, PTSD, OCD: 71–96% use waitlist controls. Published effect sizes for these disorders are likely inflated relative to care-as-usual comparisons.
-
Depression: Only 42% use waitlist, 55% use care-as-usual. Depression effect sizes are less inflated by control-condition heterogeneity than anxiety-disorder effect sizes.
-
BPD: 100% use care-as-usual. Effect sizes for BPD are not inflated by waitlist comparisons.
-
Pill placebo is rare: Only 5% of trials overall use pill placebo. This means psychotherapy efficacy is rarely compared against the same control condition used in psychiatric medication trials, making direct medication-psychotherapy comparisons methodologically fraught.
1.3 Why waitlist suppresses spontaneous recovery
Several mechanisms have been proposed:
"Waiting to change" until treatment begins. Patients assigned to waitlist may consciously or unconsciously defer problem-solving efforts until they receive the promised intervention, reducing the likelihood of spontaneous recovery during the waiting period.
Expectancy asymmetry. Patients assigned to waitlist know they are not yet receiving treatment, while patients in care-as-usual may receive supportive interventions (e.g., general practitioner visits, medication, informal counseling) that carry positive expectancies.
Disappointment effects. Participants randomized to waitlist after consenting to a trial may experience disappointment at not receiving immediate treatment, potentially worsening symptoms or reducing motivation for self-initiated improvement.
Differential attrition. Patients who improve spontaneously may drop out of waitlist conditions at higher rates than from active-treatment or care-as-usual conditions, creating selection bias in completers analyses.
None of these mechanisms are well-established. The critical empirical fact is that waitlist conditions produce less pre-post improvement (g = 0.37) than care-as-usual conditions (g = 0.64), independent of the psychotherapy received by the active-treatment arm.
Part II — Absolute outcomes across disorders
2.1 Response rates for psychotherapy across eight disorders
Cuijpers et al. (2024, World Psychiatry) pooled response rates (≥50% symptom reduction from baseline to post-test) for 441 trials across eight disorders.
Response rates in psychotherapy conditions:
| Disorder | Response rate | 95% CI | I² (%) | n (studies) |
|---|---|---|---|---|
| Depression | 0.42 | 0.39–0.45 | 82 | 159 |
| PTSD | 0.38 | 0.33–0.43 | 74 | 69 |
| OCD | 0.38 | 0.30–0.47 | 65 | 22 |
| Panic disorder | 0.38 | 0.33–0.43 | 77 | 48 |
| GAD | 0.36 | 0.30–0.42 | 74 | 48 |
| Social anxiety | 0.32 | 0.29–0.37 | 66 | 52 |
| Specific phobia | 0.32 | 0.23–0.42 | 78 | 22 |
| BPD | 0.24 | 0.15–0.36 | 82 | 21 |
Response rates in control conditions:
| Disorder | Response rate | 95% CI | I² (%) |
|---|---|---|---|
| Depression | 0.19 | 0.17–0.21 | 72 |
| PTSD | 0.10 | 0.08–0.13 | 44 |
| Panic disorder | 0.16 | 0.13–0.20 | 61 |
| GAD | 0.15 | 0.11–0.19 | 71 |
| Social anxiety | 0.12 | 0.09–0.14 | 45 |
| Specific phobia | 0.09 | 0.06–0.12 | 0 |
| OCD | 0.05 | 0.03–0.07 | 0 |
| BPD | 0.15 | 0.10–0.21 | 75 |
Interpretation:
-
Modest absolute response rates. Across all disorders, fewer than half of psychotherapy recipients achieve ≥50% symptom reduction. The highest response rate is 0.42 (depression); the lowest is 0.24 (BPD). Most patients receiving psychotherapy do not respond.
-
Control-condition response rates are non-trivial for depression. Depression control conditions produce 0.19 response—nearly one in five control participants achieves 50% symptom reduction without the experimental psychotherapy. This is consistent with known rates of spontaneous recovery and placebo response in depression.
-
Control-condition response rates are near-zero for OCD and specific phobia. OCD control conditions produce 0.05 response; specific phobia 0.09. These disorders show minimal spontaneous recovery during trial follow-up periods (typically 12–16 weeks).
-
Relative risks obscure modest absolute effects. Depression produces RR 2.09 (therapy 0.42 versus control 0.19). This is reported as "more than twice as likely to respond." But absolute response rates are 42% versus 19%—a 23-percentage-point difference. Five patients must be treated for one additional response (NNT = 4.8).
2.2 Remission versus response
Casacalenda et al. (2002) compared remission rates (achieving a score below a predetermined cutoff, typically HRSD ≤6 or ≤7) rather than response rates (≥50% reduction).
In one included trial (Elkin et al.), significantly fewer patients achieved remission (84 of 239, 35.1%) than response (116 of 239, 48.5%; p = 0.004). This pattern—response exceeding remission—indicates that many "responders" retain residual symptoms.
Translation: Response (50% reduction) is a lower bar than remission (minimal symptoms). Reporting response rates inflates apparent treatment success relative to reporting remission rates. Both metrics are clinically meaningful, but they measure different outcomes.
Part III — Psychotherapy versus medication: direct comparisons
3.1 The Casacalenda meta-analysis: nearly identical remission rates
Casacalenda et al. (2002, American Journal of Psychiatry) identified six randomized trials comparing antidepressant medication, psychotherapy, and control conditions in outpatients with major depressive disorder. The trials included 883 patients (261 medication, 352 psychotherapy, 270 controls) with mild-to-moderate, primarily nonmelancholic, nonpsychotic depression. Treatment duration ranged from 10 to 34 weeks (median 16 weeks).
Medications: Tricyclic antidepressants (amitriptyline, imipramine, nortriptyline) and phenelzine (MAO inhibitor).
Psychotherapies: Cognitive behavior therapy (3 trials), interpersonal therapy (3 trials), problem-solving therapy (1 trial), social work counseling (1 trial). All delivered in individual format.
Control conditions (critical detail):
- Pill placebo plus clinical management (3 trials)
- Treatment-as-usual by a general physician (2 trials; ~45% of participants received antidepressant medications)
- Low-intensity supportive therapy (1 trial)
Intent-to-treat remission rates:
- Medication: 46.4% (6 treatment groups)
- Psychotherapy: 46.3% (8 treatment groups)
- Control conditions: 24.4% (6 treatment groups)
- Statistical test: χ² = 37.52, df = 2, p < 0.0001
Medication and psychotherapy were both significantly more efficacious than control conditions (p < 0.0001), but were not significantly different from each other.
Completers analysis:
- Medication: 61.3% (46 of 75, 3 treatment groups)
- Psychotherapy: 58.5% (79 of 135, 4 treatment groups)
- Control conditions: 26.2% (32 of 122, 3 treatment groups)
- Statistical test: χ² = 34.47, df = 2, p < 0.0001
Sensitivity analyses: Excluding individual studies with unusually high dropout rates or atypical samples (e.g., all patients with atypical depression) did not change the findings.
Meta-analytic approach (treating remission percentages as continuous variables):
- Psychotherapy: 47.9% (95% CI: 37.8–57.9%)
- Medication: 46.2% (95% CI: 37.6–54.8%)
- Control conditions: 27.7% (95% CI: 15.7–39.7%)
- F-statistic: F = 6.82, df = 2, 17, p = 0.007
Critical point about the 24.4% control rate: Casacalenda 2002 reports 24.4% as the pooled mixed-control remission rate, not as waitlist. The control conditions were pill placebo plus clinical management, treatment-as-usual by a general physician (with about 45% receiving antidepressants), and low-intensity supportive therapy. The paper does not split remission by control type.
3.2 Dropout rates favor psychotherapy
Significantly more patients dropped out of control conditions (54.4%, 80 of 147) than medication conditions (37.1%, 46 of 124) or psychotherapy conditions (22.2%, 45 of 203; χ² = 38.54, df = 2, p < 0.0001).
Translation: Psychotherapy had the lowest dropout rate (22%), medication intermediate (37%), controls highest (54%). Lower dropout may reflect greater acceptability, fewer intolerable side effects, or stronger therapeutic alliance. It may also reflect greater efficacy—patients who experience symptom relief are less likely to drop out.
3.3 Generalizability limits
The Casacalenda trials studied:
- Outpatients with mild-to-moderate, primarily nonmelancholic, nonpsychotic depression
- Older-generation antidepressants (tricyclics, phenelzine) rather than SSRIs or SNRIs
- Treatment durations of 10–34 weeks (median 16 weeks)
Results may not generalize to:
- Severe depression, psychotic depression, treatment-resistant depression
- Bipolar depression
- Comorbid substance use disorders or serious medical conditions
- Newer antidepressants (SSRIs, SNRIs, bupropion, mirtazapine)
- Routine clinical practice, where ineffective medications are switched after 4–6 weeks rather than continued for 16 weeks
The head-to-head equivalence is established for the specific population and medications tested. Extrapolation beyond that requires caution.
Part IV — The dodo bird verdict and common factors
4.1 The common-factors model
Wampold (2015, World Psychiatry) outlines the contextual model of psychotherapy, which posits three pathways through which psychotherapy produces benefits:
Pathway 1: The real relationship. A confidential, non-judgmental connection with an empathic and caring individual, providing belongingness and reducing loneliness. Humans are social species; perceived loneliness is a mortality risk factor equal to or exceeding smoking and obesity.
Pathway 2: Expectations. Psychotherapy provides an adaptive explanation for the patient's difficulties (replacing maladaptive "folk psychology") and creates the expectation that participating in therapy will lead to improvement. Agreement about goals and tasks (the therapeutic alliance) is critical to this pathway.
Pathway 3: Specific ingredients (health-promoting actions). Every cogent treatment elicits healthy patient actions: thinking less maladaptively (CBT), improving interpersonal relations (IPT, dynamic therapy), accepting oneself (ACT, self-compassion), expressing difficult emotions (emotion-focused therapy), taking others' perspectives (mentalization). The contextual model posits that these actions are universally salubrious, not specific remedies for particular psychological deficits.
The dodo-bird prediction: If all bona fide psychotherapies (those with a cogent rationale, agreed-upon tasks, and therapist belief in the treatment) elicit health-promoting actions, then all should produce approximately equal effects. The name comes from Alice in Wonderland: "Everybody has won, and all must have prizes."
4.2 Evidence for common factors
Wampold (2015) meta-analyzes the alliance-outcome literature and related constructs:
| Common factor | Effect size (d) | n (studies) | N (patients) |
|---|---|---|---|
| Alliance | 0.57 | ~200 | >14,000 |
| Goal consensus/collaboration | 0.72 | 15 | — |
| Empathy | 0.63 | 59 | — |
| Positive regard/affirmation | 0.56 | 18 | — |
| Congruence/genuineness | 0.49 | 18 | — |
Alliance is the most researched common factor. Typically measured at session 3 or 4 and correlated with final outcome, the alliance-outcome correlation is r = 0.27 (equivalent to d = 0.57). Critics argue this could be due to:
-
Early symptom relief causes strong alliance (early responders report better alliances and have better outcomes). However, studies controlling for early progress or using longitudinal designs find that alliance predicts future symptom change even after controlling for already-occurring change.
-
Patient characteristics drive both alliance and outcome (some patients form strong alliances and also have better prognoses). However, multilevel analyses (Baldwin et al., cited in Wampold 2015) show that therapist contributions to alliance predict outcome, while patient contributions do not. More effective therapists form strong alliances across a range of patients, including those with poor attachment histories.
-
Halo effects (patients who rate alliance highly also rate outcomes highly). However, the alliance-outcome association remains robust when alliance and outcome are rated by different people.
Empathy, positive regard, and genuineness show medium-to-large effects (d = 0.49–0.63). An experimental demonstration: in a sham-acupuncture trial for irritable bowel syndrome, an "augmented relationship" condition (warm, empathic conversation about symptoms, lifestyle, and meaning of the disorder) produced superior outcomes to a "limited interaction" condition (brief, non-conversational contact), even though both received sham acupuncture. Both were superior to treatment-as-usual (waitlist).
4.3 The contested claim: are all bona fide therapies equivalent?
Wampold's position: Meta-analyses comparing bona fide therapies (excluding studies of treatments without established efficacy or inadequate therapist training) show minimal differences between therapy types. When differences are found, they are often attributable to researcher allegiance (investigators who prefer a particular therapy obtain larger effects for that therapy).
Critics' position:
-
Statistical power. Many meta-analyses pool small trials and conclude "no significant difference." Non-significance is not evidence of equivalence; it may reflect insufficient power to detect real differences.
-
Heterogeneous inclusions. Meta-analyses labeled "bona fide" often include therapies that lack evidence for the disorder being treated. For example, including psychodynamic therapy trials for panic disorder (where CBT is the established treatment) dilutes measured treatment differences.
-
Disorder-specific effects. Some disorders may show treatment differences while others do not. Aggregating across disorders obscures this heterogeneity.
4.4 Evidence for treatment differences (from Cuijpers et al., 2024)
Subgroup analyses in the 441-trial meta-analysis found:
Significant response-rate differences by therapy type:
- Depression: p < 0.001
- Panic disorder: p = 0.02
- Specific phobia: p = 0.003
- PTSD: p < 0.001
No significant response-rate differences by therapy type:
- Social anxiety disorder: p = 0.12
- GAD: p = 0.31
- OCD: p = 0.65
- BPD: p = 0.59
Translation: Treatment-type differences exist for some disorders (depression, panic, specific phobia, PTSD) but not others (social anxiety, GAD, OCD, BPD). The dodo-bird verdict is partially true: for some disorders, all bona fide therapies produce similar outcomes; for others, specific therapies outperform others.
The Cuijpers findings do not resolve the debate, because:
- They do not report which therapies are more effective for depression, panic, specific phobia, or PTSD—only that significant differences exist.
- They do not control for researcher allegiance.
- They pool waitlist, care-as-usual, and pill-placebo controls, which may differentially inflate effects for particular therapies.
Part V — What the two overclaims get wrong
5.1 The optimistic overclaim: "Therapy has huge effects"
The claim: Psychotherapy produces large effect sizes (Cohen's d > 0.8) for most mental disorders, and most patients who receive psychotherapy recover.
What's wrong with it:
-
Effect sizes are inflated by waitlist controls. Depression versus waitlist: g = 0.95. Depression versus care-as-usual: g = 0.63. The larger figure is typically cited.
-
Effect sizes are not response rates. g = 0.63 sounds impressive. But the corresponding response rates are 0.42 (therapy) versus 0.19 (controls)—a 23-percentage-point difference. Fifty-eight percent of therapy recipients do not achieve 50% symptom reduction.
-
Relative risks magnify modest effects. "More than twice as likely to respond" (RR 2.09) sounds large. But it compares 42% to 19%, not 80% to 40%.
-
Publication bias and selective reporting inflate estimates. For antidepressants, publication bias inflated effect sizes by 32% (Turner et al., 2008, discussed in the placebo review). Psychotherapy trials may be subject to similar biases, though the Cuijpers team has not published a comparable FDA-data analysis for psychotherapy.
The kernel of truth: Psychotherapy is more effective than no treatment or waitlist for most disorders. The effects are real. But they are modest, not large, and most patients do not achieve full remission.
5.2 The dismissive overclaim: "Therapy is no better than talking to a friend"
The claim: Any warm, supportive conversation produces the same outcomes as psychotherapy. Psychotherapy is merely a professionalized version of informal social support, and paying for therapy is a waste of money.
What's wrong with it:
-
Informal support lacks the three pathways of the contextual model. Talking to a friend provides Pathway 1 (real relationship) but not Pathway 2 (cogent explanation, creation of adaptive expectations, agreement about goals and tasks) or Pathway 3 (structured health-promoting actions). The contextual model predicts—and the evidence supports—that all three pathways are necessary for psychotherapy effects.
-
Control conditions that approximate "talking to a friend" produce worse outcomes than psychotherapy. In the Casacalenda trials, the pooled control conditions (pill placebo plus clinical management, treatment-as-usual by a general physician with ~45% receiving antidepressants, and low-intensity supportive therapy) produced 24.4% remission versus 46.3% for psychotherapy. The paper does not report remission rates separately for each control type. These mixed controls—which include some active elements like physician contact and medication—still produced half the remission rate of structured, manualized psychotherapy.
-
Therapist effects are substantial and reflect skill, not just warmth. More effective therapists form strong alliances across a range of patients, including those with poor attachment histories. This is a learnable skill, not an innate personality trait. Training, supervision, and deliberate practice improve therapist effectiveness.
-
Alliance alone does not account for all therapy effects. Alliance accounts for d = 0.57. But psychotherapy versus care-as-usual produces g = 0.63. Alliance is a large component, but not the entirety, of psychotherapy effects. Specific ingredients (Pathway 3) and expectations (Pathway 2) contribute additional variance.
The kernel of truth: Common factors (alliance, empathy, expectations, therapist differences) account for a large share of psychotherapy effects. Specific therapeutic ingredients (e.g., cognitive restructuring in CBT, interpretation in psychodynamic therapy) account for less variance than proponents of disorder-specific treatments claim. But this does not mean "any warm conversation" is equivalent to structured psychotherapy delivered by a trained, supervised therapist.
Part VI — Implications for evidence interpretation
6.1 Always name the comparator
Any psychotherapy efficacy claim should specify:
- Control condition type: waitlist, care-as-usual, pill placebo, psychological placebo, or another therapy.
- Outcome metric: response (≥50% reduction), remission (below clinical threshold), effect size (SMD), or functional improvement.
- Sample characteristics: severity, chronicity, comorbidity, treatment setting (primary care versus specialty mental health).
Example of a clean claim: "CBT for panic disorder produces response rates of 0.38 (95% CI: 0.33–0.43) compared to 0.16 (95% CI: 0.13–0.21) for control conditions, which were 75% waitlist, 12.5% care-as-usual, and 12.5% pill placebo. NNT = 5.0."
Example of an unclean claim: "CBT is highly effective for panic disorder, with large effect sizes and strong evidence from randomized trials."
The second claim omits the comparator, the outcome metric, and the effect magnitude. It should be mentally adjusted downward.
6.2 Distinguish between response, remission, and recovery
Response: ≥50% symptom reduction. Many responders retain residual symptoms.
Remission: Score below a clinical threshold (e.g., HRSD ≤7). Minimal symptoms, but not necessarily sustained.
Recovery: Remission sustained for a defined period (e.g., 4–6 months). Only one of the Casacalenda trials required sustained remission; most define remission at a single time point.
Reporting response inflates apparent success relative to reporting remission or recovery.
6.3 Absolute outcomes matter more than relative risks
Relative risk: Ratio of therapy response rate to control response rate. "More than twice as likely" sounds impressive.
Absolute risk reduction: Difference in response rates. 42% versus 19% = 23 percentage points.
Number needed to treat: 1 / absolute risk reduction = 1 / 0.23 ≈ 4.8. Five patients must be treated for one additional response.
NNT is the most transparent metric for clinical decision-making. An NNT of 5 means four patients receive therapy without benefit (relative to the control condition) for every one who benefits. This does not mean therapy is ineffective—it means the effects are modest and unevenly distributed.
6.4 The dodo bird debate is about mechanism, not clinical utility
The mechanism question: Do specific therapeutic ingredients (cognitive restructuring, exposure, interpretation) produce differential effects, or do common factors (alliance, empathy, expectations) account for most variance?
The clinical utility question: Should clinicians select therapies based on disorder-specific guidelines (e.g., CBT for panic, IPT for depression with interpersonal problems), or can they deliver any bona fide therapy with equal effectiveness?
The evidence suggests:
- For some disorders (depression, panic, specific phobia, PTSD): treatment differences exist. Disorder-specific selection is justified.
- For other disorders (social anxiety, GAD, OCD, BPD): no treatment differences detected in meta-analysis. Common factors may dominate.
- Therapist effects are large: delivering a therapy the therapist believes in, with strong alliance-building skill, may matter more than which specific therapy is chosen.
Part VII — What remains unresolved
7.1 Which patients benefit and which do not
Response rates of 24–42% mean that 58–76% of patients do not achieve 50% symptom reduction. Predictors of treatment response are not well-established. Baseline severity, chronicity, comorbidity, and personality disorders predict overall prognosis but do not reliably predict differential response to psychotherapy versus medication or to one therapy versus another.
Implication: Clinicians cannot reliably predict which patients will benefit from psychotherapy at baseline. Sequential treatment (trying psychotherapy, then medication, then combined treatment) is common in practice but poorly studied.
7.2 The role of specific ingredients
Treatment-type differences exist for some disorders but not others. This suggests that specific ingredients matter sometimes. But the Cuijpers analyses do not identify which therapies are superior for depression, panic, specific phobia, or PTSD. Component analyses (comparing full therapy to therapy with specific ingredients removed) show inconsistent results: some find that specific ingredients contribute above and beyond common factors; others do not.
Implication: The relative importance of specific ingredients versus common factors remains contested. Both likely contribute, but their relative weights vary by disorder, therapy type, and patient characteristics.
7.3 Long-term outcomes and relapse prevention
Most trials report outcomes at 12–16 weeks post-treatment. Relapse rates, recovery rates, and long-term functional outcomes are less well-studied. Some evidence suggests that psychotherapy (particularly CBT) produces more durable effects than medication, because patients learn skills that persist after treatment ends. But this is not consistently replicated, and long-term follow-up studies are vulnerable to attrition bias.
Implication: Short-term response rates may overestimate long-term benefit if relapse rates are high, or underestimate long-term benefit if psychotherapy produces delayed effects or skill-building that prevents future episodes.
7.4 Publication bias and selective reporting
The Cuijpers team adjusts for publication bias using Duval and Tweedie's trim-and-fill procedure, which had minimal effects for most disorders. But this method assumes that unpublished trials differ only in effect size, not in quality or design. For antidepressants, FDA data revealed that 31% of trials were not published and 11 additional trials were published as positive despite negative FDA conclusions (Turner et al., 2008). No comparable FDA-data analysis exists for psychotherapy.
Implication: Published psychotherapy effect sizes may be inflated by selective reporting, but the magnitude of inflation is unknown.
Conclusion
Psychotherapy produces modest, real benefits for most mental disorders. Response rates range from 24% to 42% across disorders; control-condition response rates range from 5% to 19%. Psychotherapy is significantly more effective than waitlist, care-as-usual, and (in the few trials that use it) pill placebo.
The effect size you measure depends on what you compare it to. Psychotherapy versus waitlist produces g = 0.95 for depression; psychotherapy versus care-as-usual produces g = 0.63. The inflation is due to reduced spontaneous recovery in waitlist, not enhanced therapy response. Any psychotherapy efficacy claim that does not name the comparator should be mentally adjusted downward.
Psychotherapy and medication produce nearly identical remission rates in direct head-to-head trials of outpatients with mild-to-moderate depression (46.4% medication, 46.3% psychotherapy, 24.4% mixed controls). Psychotherapy has lower dropout rates (22%) than medication (37%) or controls (54%).
Common factors (alliance, empathy, expectations, therapist differences) account for a large share of psychotherapy effects. Alliance produces d = 0.57; empathy d = 0.63. Therapist contributions to alliance predict outcome; patient contributions do not. More effective therapists form strong alliances across a range of patients.
Treatment-type differences exist for some disorders (depression, panic, specific phobia, PTSD) but not others (social anxiety, GAD, OCD, BPD). The dodo-bird verdict—that all bona fide therapies produce equivalent outcomes—is partially true. It depends on the disorder.
Neither "therapy has huge effects" nor "therapy is just talking to a friend" survives contact with the evidence. Psychotherapy produces real but modest effects. It requires structure, alliance-building, a cogent explanation, agreement about goals and tasks, and health-promoting actions. Informal support lacks these elements and produces smaller effects. But specific therapeutic ingredients account for less variance than disorder-specific treatment guidelines imply.
The most useful takeaway: Psychotherapy works, modestly, for most people, most of the time, compared to no treatment or waitlist. It works about as well as medication for mild-to-moderate depression. It is not a panacea. Most patients do not achieve full remission. The control condition determines the measured effect size. Always name the comparator.
About the author
Paul Stephen
Founder, Apatheia Labs
Evidence-governed research publication — Prosoche applied in the open.
Read next
More in Evidence- EssayAugust 2026
Adult ADHD and the Evidence
Prospective cohort studies find that childhood ADHD persistence to adulthood is low across named cohorts (Dunedin 5%, Pelotas 17.2%, E-Risk 21.9%), most adult-ADHD cases lack prospectively diagnosed childhood ADHD (Dunedin 90%, Pelotas 84.6%, E-Risk 67.5%), and after excluding substance use and other disorders, fewer than 1% of originally undiagnosed children meet late-onset criteria.
- EssayAugust 2026
ADHD Medications and the Evidence
Guanfacine and clonidine show child clinician SMDs −0.67 and −0.71 (Cortese 2018), but atomoxetine and viloxazine carry boxed warnings for suicidal ideation. Adult lisdexamfetamine Castells 2018 SMD −1.06; bupropion Verbeeck 2017 pooled −0.50 but clinician-only ns. Adult ER-MPH Boesen 2022 investigator SMD −0.42; IR-MPH Cândido 2021 investigator MD −20.70 but participant ns. Adult GXR Iwanami 2020 Japan diff −4.28; child GXR Sallee 2009 LS-mean −5.41 to −7.88.
- EssayAugust 2026
Antipsychotics and the Evidence
All antipsychotics beat placebo in acute trials (SMD 0.33–0.88), but 74% of CATIE participants discontinued within 18 months, dose reduction in first-episode patients doubled 7-year recovery rates versus maintenance treatment, and long-term observational data show better outcomes off medication—confounded by selection, but challenging the 'indefinite treatment for all' assumption.