🧘 Stress · 11 min read · Subtopic 2 of 5

The Effect Sizes, Honestly

Wellness content talks about journaling as if twenty minutes with a notebook reorganizes your immune system. The research record says something quieter and more interesting: across every randomized trial researchers can find, expressive writing produces a real but small effect — roughly a standardized mean difference of 0.15 — and that number has been shrinking as the evidence base grows. This page walks through where the famous bigger numbers came from, why they deflated, and the moderator patterns that predict when writing helps more or less.

🔎 Evidence Snapshot ★★★☆☆ Many trials, small pooled effects — robust existence, modest size

What the evidence supports

  • Expressive writing beats neutral-writing controls across 146 randomized trials; the pooled effect is small (r = .075, about d = 0.15) but statistically robust (Frattaroli, 2006).
  • Effects run larger for people managing medical conditions or trauma histories than for healthy college samples.
  • Benefits concentrate in physical-health reports and subjective impact; effects on psychological distress are smaller and noisier (Frisina et al., 2004).

What remains uncertain

  • Why the effect exists — inhibition and cognitive-processing accounts remain indirect.
  • How long benefits last; most trials measure weeks, not years.
  • How much expectation and publication biases inflate the published picture.

Evidence last reviewed: September 17, 2026. Conclusions may change as new research is published.

small effects, closely measured

The First Estimate Was the Loudest

The founding study — Pennebaker and Beall's 1986 four-day writing experiment, covered in the Pennebaker paradigm — was followed by a decade of replications, and in 1998 Joshua Smyth pooled thirteen of them into the field's first meta-analysis. The result: a mean effect size of d = 0.47, a medium-sized effect by psychology's conventions. Smyth illustrated it with a binomial display — illness rates of roughly 61 percent in controls versus 38 percent in writers — and that framing traveled far. It is still the number behind most enthusiastic claims you will meet online.

Two structural facts about that 1998 pool deserve attention before trusting it as "the" effect. It used a fixed-effects model, which assumes one true effect shared by all studies — an assumption the heterogeneity of this literature does not support. And most of the thirteen trials came from the founding research group, with only unpublished work underrepresented. A first meta-analysis of a young field is often an optimistic reading, and this one was.

The Evidence Grew Up and the Number Shrank

Two larger syntheses followed, and both pulled downward. Frisina, Borod, and Lepore (2004) restricted their pool to nine trials in clinical populations — people with diagnosed medical or psychiatric conditions — and found an overall effect of d = 0.19, with an instructive split: physical-health outcomes showed d = 0.21 while psychological outcomes showed d = 0.07, which was not statistically distinguishable from zero. Then Frattaroli (2006) cast the widest net: 146 randomized studies, nearly 11,000 participants, half of them unpublished dissertations and reports chased down specifically to correct publication bias, analyzed under a random-effects model. The pooled effect was r = .075 — Cohen's d ≈ 0.15, with a confidence interval of .051 to .098 that excludes zero. When she restricted the pool to higher-quality trials, the estimate eased to about d = 0.13 and stayed significant. The effect is real; it is also about a third the size the field first announced.

Four pooled estimates, one shrinking story
Standardized mean differences (Cohen's d) from the three landmark meta-analyses of expressive-writing trials. Wider nets, unpublished studies, and stricter quality screens all move the estimate down — from 0.47 to the 0.13–0.15 range.
Smyth 1998, 13 trials d = 0.47 Frisina 2004, 9 trials d = 0.19 Frattaroli 2006, 146 trials d = 0.15 Frattaroli 2006, best trials d = 0.13

What d = 0.15 Means in Human Units

Effect sizes are unreadable without translation, so here are three honest translations. As a percentile shift: an effect of d = 0.15 moves the average person from the 50th to roughly the 56th percentile on the outcome measured — visible in aggregate, easy to miss in any individual, including yourself. As a contrast: classic psychotherapy meta-analyses place talk therapy near d = 0.8 (Smith, Glass, and Miller, 1980), several times larger, though against a different comparison and a different cost. As a cost-benefit ledger: the writing intervention asks for a pen and fifteen to twenty minutes on a few days. A small effect attached to a nearly free, side-effect-light practice can still clear a rational bar — which is a different claim than "transformative," and the one the data support.

d = 0.15pooled effect across all 146 randomized trials (r = .075)
56thpercentile where the average participant lands, starting from the 50th
1/3the size of the first meta-analytic estimate (0.47 → 0.15)

The Moderator Map: When Writing Works Better

The pooled number is an average over wildly different studies, and Frattaroli's moderator analyses — cross-checked against the earlier syntheses — sketch where the effect concentrates. These are patterns across trials, not personal guarantees; they tell you which conditions the effect has historically liked. The minimum dose page turns the dose findings into a practical schedule, and the technique family covers the variants.

ModeratorPattern across trialsRead
🩺 Who writesStudies enrolling only people with medical conditions or trauma histories show larger effects than college-student samples (Frattaroli, 2006)Consistent
📅 Topic recencyTrials writing about recent events (months old, not years) show larger effects; recency correlated with effect size around r = .28Consistent
⏱️ Dose floorAt least three sessions and at least 15 minutes each beat shorter designs; adding more sessions, or spacing them out, added nothing detectableDose floor
🏠 SettingSessions at home and in private showed larger effects than supervised laboratory or group settingsPlausible
🚻 Gender mixSamples with more men showed larger effects in both Smyth (1998) and Frattaroli (2006) — possibly because men have fewer default outlets for disclosureMixed
🎯 ExpectationsWhen control groups were told nothing about benefits while writers expected them, effects roughly tripled (r = .191 vs .068 in equal-expectation trials)Inflation risk

Just as telling is what did not matter: writing by hand versus typing versus talking, daily versus weekly spacing, and — against intuition — whether the topic was framed negatively or positively all failed to moderate the pooled effect meaningfully. The active ingredient appears to be the structured disclosure itself, not the ritual around it.

Where the Inflation Comes From

Knowing why public claims outrun the data makes you resistant to both hype and cynicism. The gap has three named sources, all measured in the meta-analytic record:

So when a headline promises that journaling "rewires" stress physiology, you now have the inspection checklist: Does it cite a pooled estimate or one trial? Compared with what control? Measured on which outcome, over how long? The same literacy applies next door in meditation's effect-size map, where a similar pattern — real effects, modest size, enthusiastic coverage — plays out.

Watch-Items and Honest Limits

⚠️ Small effect is not no effect — and not a license to skip care

A d = 0.15 effect from twenty minutes of writing is a genuinely good return on effort — and it is not a treatment plan. If writing is attractive because insomnia, low mood, or blood-pressure spikes are wearing you down, treat those signals as the primary project with a clinician, and the notebook as one low-cost tool alongside it.

Questions, Answered Briefly

The Bottom Line

  1. The honest pooled effect is small — about d = 0.15 across 146 randomized trials (r = .075), roughly a third of the field's first 1998 estimate.
  2. The shrinkage is informative, not damning — publication bias, expectation effects, and outcome selection explain most of the gap between headlines and pooled data.
  3. Moderators are the practical content — effects run larger for medical and trauma samples, recent topics, at least three sessions of 15+ minutes, and private home settings.
  4. Price it correctly — a small, nearly free, low-risk effect can be worth keeping; it is not a substitute for clinical care when symptoms are the real project.

Related Topics

Sources & further reading