Receipts · research july 2026 · published 2026-08-03 · v1 · 4 min read
The evidence trap
Where proving a thing works is expensive and not proving it is legal, the bar taxes the honest product
What the shutdown of an FDA-designated chatbot in consumer mental health shows about markets where validation is expensive and silence is free, including the second tax nobody budgets for. The canonical treatment of the evidence trap.
The comfortable version of this argument says that evidence eventually wins, that a product which submits to measurement outlasts the ones that decline. We would like to carry that version. The record does not support it, and the clearest counter-case is recent. On June 30, 2025, Woebot Health shut down its direct-to-consumer application. The company had been founded in 2017 by a clinical research psychologist, had raised roughly $123 million, and in 2021 had earned an FDA Breakthrough Device Designation for WB001, its therapeutic for postpartum depression. Its founder told STAT that the shutdown was largely attributable to the cost and challenge of meeting the agency’s requirements for marketing authorization. One of the few products in the category to have earned a federal breakthrough designation left the consumer market, and the products that had never gathered evidence at all stayed.
The trap is built into where the bar sits. Medical-device regulation attaches to a product’s intended use, which means the evidence bar is triggered by what a product holds itself out to do rather than by the mechanism it uses. Two applications can run identical logic and behave identically on a phone, and only the one that says out loud what it is for becomes a device. Industry figures put a De Novo classification at a median around $5 million and roughly five and a half years from concept to decision. So the honest company pays that, and the company that implies the same benefit while claiming nothing pays none of it, and both compete for a user who cannot see the difference. The money spent on proof transfers, in competitive effect, to whoever declined to gather any.
Then the second tax arrives, and almost nobody budgets for it. Validation requires reproducible output, so a product built to pass it is constrained to determinism. Woebot ran on pre-scripted rules-based therapy content written by clinicians, which was not a failure of ambition but the price of admission to the pathway it was on. When fluent generators arrived, those scripts read as archaic beside them, and the founder was explicit that the company wanted to use large language models and that the agency had not yet worked out how to regulate them. The discipline that earned the evidence is the same discipline that made the product feel old. Punished at the treasury, then punished again at the interface, and neither penalty had anything to do with whether the thing worked.
The market underneath all of this is thin to begin with. Across 93 mental health applications, one systematic analysis found a median daily open rate of 4.0 percent and median 30-day retention of 3.3 percent. Into that, a regulated product carries device overhead its unregulated neighbor does not. Where the lines have been drawn since, they fall in the same place rather than a better one. Utah’s H.B. 452, in effect since May 2025, defines a regulated mental health chatbot in terms of generative technology used in conversations a reasonable person would construe as therapy, and excludes tools that deliver scripted output or hand a person to a human. Read as a fact about how regulation works rather than as a map of where to hide, it says what the federal regime says. The trigger is the claim.
This is the case where restraint lost, and a corpus that only publishes its wins is running the same selective measurement it objects to elsewhere. The conclusion is not that proving things is a mistake. It is that proof is priced in two currencies, money and datedness, and a builder who budgets only the first will discover the second at the worst possible moment. Know which taxes you have agreed to pay before you agree to them. And notice what the trap actually is, because the naming decides what can be done about it. Nothing here is a property of evidence. It is a property of a market wired so that saying less costs less, and wiring is the kind of thing that gets rebuilt.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 4
-
STAT (reporting on Woebot Health and its founder Alison Darcy) (2025). Coverage of the Woebot therapy chatbot shutdown, July 2025
The first-party attribution for why the consumer app closed: the cost and challenge of meeting the FDA's requirements for marketing authorization, made more pressing by large language models the agency had not yet worked out how to regulate.
Comment on this source -
MobiHealthNews and contemporaneous trade coverage (2025). Reporting on the Woebot Health app shutdown, April to July 2025
The shutdown date, the capital raised, and the 2021 FDA Breakthrough Device Designation for WB001.
Comment on this source -
Amit Baumel, Frederick Muench, Stav Edan, John M. Kane (2019). Objective user engagement with mental health apps, systematic search and panel-based usage analysis (Journal of Medical Internet Research)
The independent measurement of the market a regulated product was paying to enter: across 93 apps, a median daily open rate of 4.0 percent and median 30-day retention of 3.3 percent.
Comment on this source -
Utah State Legislature (2025). H.B. 452, Artificial Intelligence Amendments (2025 general session)
Cited only as a fact about where a legislature drew its line, at how a system holds itself out rather than at any measure of whether it helps. Not treated as a route around regulation.
Comment on this source
Claims and confidence 9
- verified
Woebot Health shut down its direct-to-consumer application on June 30, 2025, having raised roughly $123 million and holding an FDA Breakthrough Device Designation granted in 2021 for WB001, its postpartum depression therapeutic.
Contemporaneous trade reporting on the April 2025 announcement and the June 2025 closure, retrieved and read during the receipts wave on 2026-08-03. The capital figure is the commonly reported total of a $90 million Series B and a later $9 million investment; some internal and secondary accounts put the lifetime total at $129 million, so the figure is written as approximate.
Respond to this claim - verified
Woebot's founder attributed the shutdown largely to the cost and challenge of meeting the FDA's requirements for marketing authorization, and said the company wanted to use large language models the agency had not yet worked out how to regulate.
Reported statements to STAT, July 2025.
Respond to this claim - verified
Woebot's conversational product ran on pre-scripted, rules-based cognitive behavioral therapy content written by clinicians rather than on generative models.
Consistent across the company's own descriptions and contemporaneous reporting.
Respond to this claim - verified
Medical-device regulation attaches to a product's intended use, so the evidence bar is triggered by what a product holds itself out to do rather than by the mechanism it uses.
The intended-use basis of device classification is a settled feature of the regime, not a finding of this piece.
Respond to this claim - directional
A De Novo classification carries a median cost of roughly $5 million and a median of about 66 months from concept to agency decision.
Commonly cited industry analysis of device pathway cost and duration, reported in trade press; the primary study was not retrieved during verification, so the figures are carried as approximate.
Respond to this claim - verified
Across 93 mental health applications, the median daily open rate was 4.0 percent and median 30-day retention was 3.3 percent.
Baumel et al., Journal of Medical Internet Research 2019, panel-based usage analysis of Android apps with 10,000 or more installs. A 2019 measurement, cited as the market's shape rather than as current figures.
Respond to this claim - verified
Utah's H.B. 452, in effect since May 2025, defines a regulated mental health chatbot in terms of generative technology used in conversations a reasonable person would construe as mental health therapy, and excludes tools that provide scripted output or facilitate connection to a human therapist.
The enrolled bill and legal-practice summaries of it, read during verification. Cited as evidence about where a line was drawn, never as guidance on qualifying for an exclusion.
Respond to this claim - directional
Where validation is expensive and an unvalidated substitute is legal, the evidence bar functions as a tax on the honest product and the market can select against the thing that proved itself.
The brick's argument, built from the claims above. One well-documented case plus a regulatory structure, not a measured market-wide effect.
Respond to this claim - directional
Because validation requires reproducible output, a product built to pass it is constrained to determinism, which reads as archaic beside a fluent generator.
An inference from the reproducibility requirement and the founder's own account of the bind. Argued, not measured.
Respond to this claim