Our evidence page says what the research supports. This page says what it does not support. Both are ours.
We traced the citations behind this platform to their primary sources. There were no fabrications — but every error we found ran in the same direction: compression toward a cleaner, stronger story than the source actually supports. Those are listed below, including the one that cost us our best-sounding claim.
Winter and colleagues (BMJ Open, 2020) report a standardised mean difference of 0.52 (95% CI 0.36 to 0.67) across 22 randomized trials. Paulus & Meinken (2022) report Hedges g = 0.58 — falling to 0.51 (95% CI 0.39 to 0.62) once two outlier studies are removed. We have read both of these and confirmed the figures. Teding van Berkhout & Malouff (2016) report g = 0.63 across 18 RCTs and 1,018 participants. That paper is closed-access and we have never read it: as of 2026-08-20 we tried Unpaywall, Semantic Scholar, PubMed Central, Crossref TDM and a publisher API, and all five failed. The figure comes from the publisher's abstract. An earlier version of this page said we were relying on a colleague's verification; that was not accurate, and no such verification exists.
Heterogeneity is substantial across the board — I² = 63% in Winter, 76.9% in Paulus & Meinken. Winter's authors rate the overall quality of evidence as LOW under GRADE, with 15 of 26 included trials at unclear or high risk of bias. "Empathy training works" does not license "our intervention works." A pooled estimate drawn from heterogeneous, low-quality trials justifies testing what we built; it does not substitute for it.
They decay. Winter's sustainability analysis (11 trials) finds 0.69 (95% CI 0.23 to 1.15) up to twelve weeks and 0.34 (95% CI 0.11 to 0.57) at twelve weeks and beyond — note how wide the first interval is. On publication bias the evidence is mixed rather than settled: Teding van Berkhout's g = 0.63 falls to 0.51 under trim-and-fill correction — a figure we have taken from that paper's abstract, not read in its text — while Winter and Fragkos & Crampton both examined it and found none. We report both because we can only honestly report both.
Paulus & Meinken (International Journal of Medical Education, 2022) coded hours of training across their included studies and report that training hours were not statistically significantly associated with effect size. We have read that paper and confirmed the finding. It is why this platform is built around one drill done well rather than a large content library: surface area appears to buy nothing.
We checked the two studies most often cited for this. Neither measured cognitive empathy with an objective instrument. One used a perspective-taking questionnaire in a design with no demonstrated baseline equivalence; the other used two self-report affective items in a survey experiment, and reported a knowledge penalty alongside the empathy gain. The gap is wider than the secondary literature suggests. That gap is what we are trying to fill. To be precise about what is missing: a pedagogical literature on debate and perspective-taking does exist — Zorwick (2016), a chapter in Using Debate in the Classroom (Routledge), is written on exactly this topic. What does not exist is a randomized trial with an objective measure. The honest claim is "no randomized trial," not "no literature."
We have not yet shown that this platform changes anything. A preregistration is drafted and not submitted, and the study it describes has not begun. Until it does, the honest claim is that the premise is supported by the literature and our specific intervention is untested.
The benchmark figures behind our misperception drill are being traced to their original survey instruments. Until each is sourced — survey house, exact question wording, field dates, sample — the app does not show a participant an unverified figure, and no analysis will use one.