The Empathy Enigma

What we checked,
and what we got wrong.

Our evidence page says what the research supports. This page says what it does not support. Both are ours.

We traced the citations behind this platform to their primary sources. There were no fabrications — but every error we found ran in the same direction: compression toward a cleaner, stronger story than the source actually supports. Those are listed below, including the one that cost us our best-sounding claim.

Claim by claim
Empathy training produces a moderate effect.
Holds, with caveats

Winter and colleagues (BMJ Open, 2020) report a standardised mean difference of 0.52 (95% CI 0.36 to 0.67) across 22 randomized trials. Paulus & Meinken (2022) report Hedges g = 0.58 — falling to 0.51 (95% CI 0.39 to 0.62) once two outlier studies are removed. We have read both of these and confirmed the figures. Teding van Berkhout & Malouff (2016) report g = 0.63 across 18 RCTs and 1,018 participants. That paper is closed-access and we have never read it: as of 2026-08-20 we tried Unpaywall, Semantic Scholar, PubMed Central, Crossref TDM and a publisher API, and all five failed. The figure comes from the publisher's abstract. An earlier version of this page said we were relying on a colleague's verification; that was not accurate, and no such verification exists.

That effect transfers to this platform.
Not established

Heterogeneity is substantial across the board — I² = 63% in Winter, 76.9% in Paulus & Meinken. Winter's authors rate the overall quality of evidence as LOW under GRADE, with 15 of 26 included trials at unclear or high risk of bias. "Empathy training works" does not license "our intervention works." A pooled estimate drawn from heterogeneous, low-quality trials justifies testing what we built; it does not substitute for it.

Effects last.
Qualified

They decay. Winter's sustainability analysis (11 trials) finds 0.69 (95% CI 0.23 to 1.15) up to twelve weeks and 0.34 (95% CI 0.11 to 0.57) at twelve weeks and beyond — note how wide the first interval is. On publication bias the evidence is mixed rather than settled: Teding van Berkhout's g = 0.63 falls to 0.51 under trim-and-fill correction — a figure we have taken from that paper's abstract, not read in its text — while Winter and Fragkos & Crampton both examined it and found none. We report both because we can only honestly report both.

More practice produces more improvement.
Unverified by us

Paulus & Meinken (International Journal of Medical Education, 2022) coded hours of training across their included studies and report that training hours were not statistically significantly associated with effect size. We have read that paper and confirmed the finding. It is why this platform is built around one drill done well rather than a large content library: surface area appears to buy nothing.

Debate specifically has been shown to build cognitive empathy.
The open gap

We checked the two studies most often cited for this. Neither measured cognitive empathy with an objective instrument. One used a perspective-taking questionnaire in a design with no demonstrated baseline equivalence; the other used two self-report affective items in a survey experiment, and reported a knowledge penalty alongside the empathy gain. The gap is wider than the secondary literature suggests. That gap is what we are trying to fill. To be precise about what is missing: a pedagogical literature on debate and perspective-taking does exist — Zorwick (2016), a chapter in Using Debate in the Classroom (Routledge), is written on exactly this topic. What does not exist is a randomized trial with an objective measure. The honest claim is "no randomized trial," not "no literature."

Corrections we made to our own material
  • A meta-analysis we rely on is dated 2016, not 2015, and carries a published correction notice. Its pooled estimate also excludes one high-effect outlier study — defensible, but worth stating.
  • The debate study we had treated as our strongest empirical anchor is a four-page piece in the practitioner section of a political science journal, and its two headline findings come from different samples. We no longer derive an effect-size anchor from it.
  • A deliberation study we cited for empathy and knowledge rising together in fact reports a trade-off: opposing viewpoints raised empathy but dampened knowledge gains.
  • A PNAS finding on out-group empathy was attributed to the wrong authors and year on our evidence page. It is Hein, Engelmann, Vollberg & Tobler, 2015. Corrected.
  • Two citations had the wrong author list entirely. Findings were extracted correctly; attribution was not.
  • Four claims on our evidence page rested on author-year citations with no DOI. Traced 2026-08-20: one was a misattribution (the finding credited to Batson et al., 1988 is Batson & Weeks, 1996), two overstated what their sources support and have been rewritten, and one — a well-being claim credited to a "Konrath, 2013 meta-review" — has been REMOVED from the page entirely. The work behind it is a book chapter with no DOI that we could not obtain, so we could not check it. A claim we cannot check does not belong on this page.
  • This page previously said a paywalled meta-analysis had been verified by a colleague. It had not been verified by anyone on our side. Corrected, and the two figures drawn from it are now labelled as abstract-only.

What we have not measured yet

We have not yet shown that this platform changes anything. A preregistration is drafted and not submitted, and the study it describes has not begun. Until it does, the honest claim is that the premise is supported by the literature and our specific intervention is untested.

The benchmark figures behind our misperception drill are being traced to their original survey instruments. Until each is sourced — survey house, exact question wording, field dates, sample — the app does not show a participant an unverified figure, and no analysis will use one.