brick · v2 · 2026-08-03
The empathy paradox
A simulation outscored doctors on empathy because the human baseline had been degraded, not because the machine feels anything
The measured case in which machine answers beat physicians on empathy, why the honest conclusion is about the physicians' working conditions, and what remains scarce once the writing is commoditized. The canonical treatment of the empathy paradox.
The professional defense against machine intelligence in the caring trades has always been the same sentence. Machines cannot care. In April 2023 a research team led by John Ayers put the sentence in front of a measurement and it did not survive contact. They took 195 patient questions posted to a public medical forum in October 2022, each already answered by a verified physician, and asked a chatbot the same questions. Three licensed health professionals then read both answers blind. Across 585 evaluations they preferred the machine’s answer 78.6 percent of the time, rated it good or very good quality in 78.5 percent of cases against 22.1 percent for the doctors, and rated it empathetic or very empathetic 45.1 percent of the time against 4.6 percent.
Almost every retelling of that result makes the same error, which is to treat it as news about the machine. It is not. Look at what was actually on the two sides of the comparison. The physicians’ replies averaged 52 words and the chatbot’s averaged 211, which is not a difference in compassion but a difference in available minutes. These were unpaid answers typed into a forum by doctors whose paid work runs on appointment slots measured in a quarter of an hour, in a profession with a documented burnout problem that predates the technology by a decade. The other side had unlimited time, no previous patient running late, no inbox, and no bad day. The comparison was never warmth against simulation. It was a writer with infinite patience against a professional with none left, and the finding grades the conditions the human reply was written under rather than anything about the machine’s interior.
That reading is less flattering to everyone and considerably more useful. It says the human baseline in these professions has been degraded to the point where a simulation of unhurried attention beats the real thing on a written page, and it locates the failure in the scheduling, the documentation load, and the economics that produced a fifteen-minute encounter, none of which are laws of nature. It also explains why the result feels wrong to clinicians who read it. They know what they are capable of when they have the time. The study did not measure that, because the study measured text.
Which is where the boundary of the finding sits, stated as flatly as the finding itself. What was rated was writing, by a panel of professionals, on a screen. Not patients, not outcomes, not anything that happened in a room between two people. The machine won the part of medicine that can be typed. Everything the corpus argues is scarce lives in the part that cannot be, which is the settled nervous system, the person who is answerable, and the bond that the psychotherapy literature keeps finding is the strongest robust predictor of whether helping work helps. When the writing is commoditized, that presence is not a soft benefit hanging off the service. It is the remaining product, and its price goes up rather than down.
So the useful response to a result like this is not the reflex on either side. The defensive reflex says the ratings must be measuring something shallow. The credulous reflex says machines are more compassionate than doctors now. Both skip the actual news, which is that we have built a system of care in which fifteen minutes of a person’s attention has become scarcer than an unlimited amount of a machine’s. Take the machine, then, for the drafting and the inbox and the long careful answer at midnight. Take the hours it gives back and put them where the measurement could not reach, because a profession that wins that comparison honestly, with the time restored and the human unhurried, has something no capability curve is coming for.