Skip to content

The Face · research january to may 2026 · published 2026-08-03 · v3 · 3 min read · history

The empathy paradox

A simulation outscored doctors on empathy because the human baseline had been degraded, not because the machine feels anything

The measured case in which machine answers beat physicians on empathy, why the honest conclusion is about the physicians' working conditions, and what remains scarce once the writing is commoditized. The canonical treatment of the empathy paradox.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "In a 2023 study of 195 patient questions from a public medical forum, a blinded panel of licensed clinicians rated chatbot responses empathetic or very empathetic 45.1 percent of the time against 4.6 percent for the physicians' own replies, and preferred the chatbot response in 78.6 percent of 585 evaluations."

    verified. Ayers et al., JAMA Internal Medicine 2023. Panel ratings of written responses by three licensed health professionals per exchange; not patient ratings and not clinical outcomes.

Open the complete evidence in the structured publication.

A blinded panel of clinicians preferred a chatbot's answers to real patient questions over the verified physicians' own replies, and rated them empathetic almost ten times as often.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "In the same study the physicians' replies averaged 52 words against the chatbot's 211, and the panel rated the chatbot's answers good or very good quality in 78.5 percent of cases against 22.1 percent for the physicians."

    verified. Ayers et al. 2023, reported means with interquartile ranges of 17 to 62 words for physicians and 168 to 245 for the chatbot.

Open the complete evidence in the structured publication.

The comparison ran between an unhurried writer with unlimited patience and a burned-out professional typing between appointments, so the result grades the conditions the human reply was written under rather than anything about the machine's interior.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Read the finding as an indictment of the baseline rather than a claim about the machine, and price what the comparison could not measure, which is the regulated human being in the room.

The professional defense against machine intelligence in the caring trades has always been the same sentence. Machines cannot care. In April 2023 a research team led by John Ayers put the sentence in front of a measurement and it did not survive contact. They took 195 patient questions posted to a public medical forum in October 2022, each already answered by a verified physician, and asked a chatbot the same questions. Three licensed health professionals then read both answers blind. Across 585 evaluations they preferred the machine’s answer 78.6 percent of the time, rated it good or very good quality in 78.5 percent of cases against 22.1 percent for the doctors, and rated it empathetic or very empathetic 45.1 percent of the time against 4.6 percent.

Almost every retelling of that result makes the same error, which is to treat it as news about the machine. It is not. Look at what was actually on the two sides of the comparison. The physicians’ replies averaged 52 words and the chatbot’s averaged 211, which is not a difference in compassion but a difference in available minutes. These were unpaid answers typed into a forum by doctors whose paid work runs on appointment slots measured in a quarter of an hour, in a profession with a documented burnout problem that predates the technology by a decade. The other side had unlimited time, no previous patient running late, no inbox, and no bad day. The comparison was never warmth against simulation. It was a writer with infinite patience against a professional with none left, and the finding grades the conditions the human reply was written under rather than anything about the machine’s interior.

That reading is less flattering to everyone and considerably more useful. It says the human baseline in these professions has been degraded to the point where a simulation of unhurried attention beats the real thing on a written page, and it locates the failure in the scheduling, the documentation load, and the economics that produced a fifteen-minute encounter, none of which are laws of nature. It also explains why the result feels wrong to clinicians who read it. They know what they are capable of when they have the time. The study did not measure that, because the study measured text.

Which is where the boundary of the finding sits, stated as flatly as the finding itself. What was rated was writing, by a panel of professionals, on a screen. Not patients, not outcomes, not anything that happened in a room between two people. The machine won the part of medicine that can be typed. Everything the corpus argues is scarce lives in the part that cannot be, which is the settled nervous system, the person who is answerable, and the bond that the psychotherapy literature keeps finding is the strongest robust predictor of whether helping work helps. When the writing is commoditized, that presence is not a soft benefit hanging off the service. It is the remaining product, and its price goes up rather than down.

So the useful response to a result like this is not the reflex on either side. The defensive reflex says the ratings must be measuring something shallow. The credulous reflex says machines are more compassionate than doctors now. Both skip the actual news, which is that we have built a system of care in which fifteen minutes of a person’s attention has become scarcer than an unlimited amount of a machine’s. Take the machine, then, for the drafting and the inbox and the long careful answer at midnight. Take the hours it gives back and put them where the measurement could not reach, because a profession that wins that comparison honestly, with the time restored and the human unhurried, has something no capability curve is coming for.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 2
  1. John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, Davey M. Smith (2023). Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum (JAMA Internal Medicine 183(6), 589-596)

    The brick's anchor. Cross-sectional study of 195 exchanges drawn from Reddit's r/AskDocs in October 2022, evaluated blind and in triplicate by licensed health professionals, 585 evaluations in total. Verified against the published paper and institutional summaries during the face wave.

    Comment on this source
  2. Bruce Wampold (2015). The Great Psychotherapy Debate (common factors research)

    The alliance finding that locates the remaining scarcity in what the study could not measure. Claim restated verbatim from the presence-dividend apparatus.

    Comment on this source
Claims and confidence 4
  1. verified

    In a 2023 study of 195 patient questions from a public medical forum, a blinded panel of licensed clinicians rated chatbot responses empathetic or very empathetic 45.1 percent of the time against 4.6 percent for the physicians' own replies, and preferred the chatbot response in 78.6 percent of 585 evaluations.

    Ayers et al., JAMA Internal Medicine 2023. Panel ratings of written responses by three licensed health professionals per exchange; not patient ratings and not clinical outcomes.

    Respond to this claim
  2. verified

    In the same study the physicians' replies averaged 52 words against the chatbot's 211, and the panel rated the chatbot's answers good or very good quality in 78.5 percent of cases against 22.1 percent for the physicians.

    Ayers et al. 2023, reported means with interquartile ranges of 17 to 62 words for physicians and 168 to 245 for the chatbot.

    Respond to this claim
  3. directional

    The measured advantage reflects the conditions the human replies were written under, unpaid forum answers by time-constrained clinicians, rather than any affective capacity in the model.

    Directional. The length and setting differences are in the paper and the burnout context is well documented, but the study did not manipulate time or workload, so the causal reading is ours and is not a result the study establishes.

    Respond to this claim
  4. verified

    The therapeutic alliance is among the strongest robust predictors of psychotherapy outcomes.

    Wampold and the common-factors literature; replicated across decades and meta-analyses.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to The empathy paradox

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The empathy paradox

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.