{
  "schema": "org-writing@v1",
  "slug": "the-evidence-trap",
  "kg": {
    "id": "org:writing:the-evidence-trap",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "The evidence trap",
  "subtitle": "Where proving a thing works is expensive and not proving it is legal, the bar taxes the honest product",
  "abstract": "What the shutdown of an FDA-designated chatbot in consumer mental health shows about markets where validation is expensive and silence is free, including the second tax nobody budgets for. The canonical treatment of the evidence trap.",
  "kind": "brick",
  "topics": [
    "Receipts"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:receipts",
      "topic": "Receipts",
      "wall": "org:walls:economics",
      "position": 4,
      "total": 6
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "A conversational product holding a federal breakthrough designation shut down its direct-to-consumer app in June 2025, after roughly $123 million, while unvalidated substitutes in the same category carried on.",
      "claims": [
        "shut down its direct-to-consumer application on June 30, 2025"
      ]
    },
    "mechanism": {
      "text": "The evidence bar is triggered by what a product holds itself out to do rather than by the mechanism it uses, so the company that proves an effect pays for the proof while the company implying the same effect pays nothing, and the market can select against the thing that proved itself.",
      "claims": [
        "triggered by what a product holds itself out to do rather than by the mechanism it uses"
      ]
    },
    "move": {
      "text": "Price both taxes before starting, the money the proof costs and the datedness the proof enforces, because an evidence trap is a property of a market's wiring rather than an argument against evidence.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-strategy",
      "path": "docs/20-business/30-competitive-analysis/defensibility-research.md"
    },
    {
      "repo": "mnstry-research",
      "path": "topics/business/strategy/projects/competitive-positioning-deep-research/research/cpr-16-woebot-shutdown-analysis.md"
    }
  ],
  "canonicalPath": "/writing/the-evidence-trap/",
  "body": "The comfortable version of this argument says that evidence eventually wins, that a product which submits to measurement outlasts the ones that decline. We would like to carry that version. The record does not support it, and the clearest counter-case is recent. On June 30, 2025, Woebot Health shut down its direct-to-consumer application. The company had been founded in 2017 by a clinical research psychologist, had raised roughly $123 million, and in 2021 had earned an FDA Breakthrough Device Designation for WB001, its therapeutic for postpartum depression. Its founder told STAT that the shutdown was largely attributable to the cost and challenge of meeting the agency's requirements for marketing authorization. One of the few products in the category to have earned a federal breakthrough designation left the consumer market, and the products that had never gathered evidence at all stayed.\n\nThe trap is built into where the bar sits. Medical-device regulation attaches to a product's intended use, which means the evidence bar is triggered by what a product holds itself out to do rather than by the mechanism it uses. Two applications can run identical logic and behave identically on a phone, and only the one that says out loud what it is for becomes a device. Industry figures put a De Novo classification at a median around $5 million and roughly five and a half years from concept to decision. So the honest company pays that, and the company that implies the same benefit while claiming nothing pays none of it, and both compete for a user who cannot see the difference. The money spent on proof transfers, in competitive effect, to whoever declined to gather any.\n\nThen the second tax arrives, and almost nobody budgets for it. Validation requires reproducible output, so a product built to pass it is constrained to determinism. Woebot ran on pre-scripted rules-based therapy content written by clinicians, which was not a failure of ambition but the price of admission to the pathway it was on. When fluent generators arrived, those scripts read as archaic beside them, and the founder was explicit that the company wanted to use large language models and that the agency had not yet worked out how to regulate them. The discipline that earned the evidence is the same discipline that made the product feel old. Punished at the treasury, then punished again at the interface, and neither penalty had anything to do with whether the thing worked.\n\nThe market underneath all of this is thin to begin with. Across 93 mental health applications, one systematic analysis found a median daily open rate of 4.0 percent and median 30-day retention of 3.3 percent. Into that, a regulated product carries device overhead its unregulated neighbor does not. Where the lines have been drawn since, they fall in the same place rather than a better one. Utah's H.B. 452, in effect since May 2025, defines a regulated mental health chatbot in terms of generative technology used in conversations a reasonable person would construe as therapy, and excludes tools that deliver scripted output or hand a person to a human. Read as a fact about how regulation works rather than as a map of where to hide, it says what the federal regime says. The trigger is the claim.\n\nThis is the case where restraint lost, and a corpus that only publishes its wins is running the same selective measurement it objects to elsewhere. The conclusion is not that proving things is a mistake. It is that proof is priced in two currencies, money and datedness, and a builder who budgets only the first will discover the second at the worst possible moment. Know which taxes you have agreed to pay before you agree to them. And notice what the trap actually is, because the naming decides what can be done about it. Nothing here is a property of evidence. It is a property of a market wired so that saying less costs less, and wiring is the kind of thing that gets rebuilt.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:the-evidence-trap:r01",
        "author": "STAT (reporting on Woebot Health and its founder Alison Darcy)",
        "work": "Coverage of the Woebot therapy chatbot shutdown, July 2025",
        "year": 2025,
        "relevance": "The first-party attribution for why the consumer app closed: the cost and challenge of meeting the FDA's requirements for marketing authorization, made more pressing by large language models the agency had not yet worked out how to regulate."
      },
      {
        "id": "org:references:the-evidence-trap:r02",
        "author": "MobiHealthNews and contemporaneous trade coverage",
        "work": "Reporting on the Woebot Health app shutdown, April to July 2025",
        "year": 2025,
        "relevance": "The shutdown date, the capital raised, and the 2021 FDA Breakthrough Device Designation for WB001."
      },
      {
        "id": "org:references:the-evidence-trap:r03",
        "author": "Amit Baumel, Frederick Muench, Stav Edan, John M. Kane",
        "work": "Objective user engagement with mental health apps, systematic search and panel-based usage analysis (Journal of Medical Internet Research)",
        "year": 2019,
        "relevance": "The independent measurement of the market a regulated product was paying to enter: across 93 apps, a median daily open rate of 4.0 percent and median 30-day retention of 3.3 percent."
      },
      {
        "id": "org:references:the-evidence-trap:r04",
        "author": "Utah State Legislature",
        "work": "H.B. 452, Artificial Intelligence Amendments (2025 general session)",
        "year": 2025,
        "relevance": "Cited only as a fact about where a legislature drew its line, at how a system holds itself out rather than at any measure of whether it helps. Not treated as a route around regulation."
      }
    ],
    "claims": [
      {
        "id": "org:claims:the-evidence-trap:c01",
        "claim": "Woebot Health shut down its direct-to-consumer application on June 30, 2025, having raised roughly $123 million and holding an FDA Breakthrough Device Designation granted in 2021 for WB001, its postpartum depression therapeutic.",
        "basis": "Contemporaneous trade reporting on the April 2025 announcement and the June 2025 closure, retrieved and read during the receipts wave on 2026-08-03. The capital figure is the commonly reported total of a $90 million Series B and a later $9 million investment; some internal and secondary accounts put the lifetime total at $129 million, so the figure is written as approximate.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c02",
        "claim": "Woebot's founder attributed the shutdown largely to the cost and challenge of meeting the FDA's requirements for marketing authorization, and said the company wanted to use large language models the agency had not yet worked out how to regulate.",
        "basis": "Reported statements to STAT, July 2025.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c03",
        "claim": "Woebot's conversational product ran on pre-scripted, rules-based cognitive behavioral therapy content written by clinicians rather than on generative models.",
        "basis": "Consistent across the company's own descriptions and contemporaneous reporting.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c04",
        "claim": "Medical-device regulation attaches to a product's intended use, so the evidence bar is triggered by what a product holds itself out to do rather than by the mechanism it uses.",
        "basis": "The intended-use basis of device classification is a settled feature of the regime, not a finding of this piece.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c05",
        "claim": "A De Novo classification carries a median cost of roughly $5 million and a median of about 66 months from concept to agency decision.",
        "basis": "Commonly cited industry analysis of device pathway cost and duration, reported in trade press; the primary study was not retrieved during verification, so the figures are carried as approximate.",
        "confidence": "directional",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c06",
        "claim": "Across 93 mental health applications, the median daily open rate was 4.0 percent and median 30-day retention was 3.3 percent.",
        "basis": "Baumel et al., Journal of Medical Internet Research 2019, panel-based usage analysis of Android apps with 10,000 or more installs. A 2019 measurement, cited as the market's shape rather than as current figures.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c07",
        "claim": "Utah's H.B. 452, in effect since May 2025, defines a regulated mental health chatbot in terms of generative technology used in conversations a reasonable person would construe as mental health therapy, and excludes tools that provide scripted output or facilitate connection to a human therapist.",
        "basis": "The enrolled bill and legal-practice summaries of it, read during verification. Cited as evidence about where a line was drawn, never as guidance on qualifying for an exclusion.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c08",
        "claim": "Where validation is expensive and an unvalidated substitute is legal, the evidence bar functions as a tax on the honest product and the market can select against the thing that proved itself.",
        "basis": "The brick's argument, built from the claims above. One well-documented case plus a regulatory structure, not a measured market-wide effect.",
        "confidence": "directional",
        "sources": []
      },
      {
        "id": "org:claims:the-evidence-trap:c09",
        "claim": "Because validation requires reproducible output, a product built to pass it is constrained to determinism, which reads as archaic beside a fluent generator.",
        "basis": "An inference from the reproducibility requirement and the founder's own account of the bind. Argued, not measured.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:evidence-trap",
        "name": "Evidence trap",
        "definition": "A market condition in which proving that a product works is expensive and offering an unvalidated substitute is legal, so the evidence bar operates as a tax on the honest product and selection runs against the thing that proved itself. The bar attaches to what a product holds itself out to do rather than to how it works, which is why silence is cheaper than proof. A second tax follows and is rarely budgeted for, since validation requires reproducible output and the resulting determinism reads as archaic beside an unvalidated fluent competitor.",
        "provenance": "canonical"
      }
    ],
    "researchContext": "Sourced from the defensibility research and the Woebot shutdown analysis. Only\nthe anatomy of the exit is used. The source documents' surrounding material,\ntheir five-failure-mode checklist framed as things a company must avoid, their\ncounter-positioning conclusions, and in particular their reading of Utah H.B.\n452 as a safe harbor to claim in investor materials, is competitive strategy\nand none of it enters the corpus. H.B. 452 appears here only as evidence about\nwhere a legislature drew its line, which is the same place the federal regime\ndraws it. Every named-company fact was re-verified by search before authoring,\nwhich is why the capital figure is written as approximate rather than carrying\nthe internal document's precise number, and why the pathway cost and duration\nare graded directional rather than verified. The two-taxes framing, the reading\nof determinism as the second and unbudgeted price of validation, and the\ninsistence that the trap is a property of market wiring rather than an argument\nagainst evidence are the brick's contribution. Published deliberately as a loss,\nsince a corpus that reports only the cases where restraint won would be running\nthe selective measurement it objects to."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "e0b6407ca3a94f5f4cf88ff0aca88bc15cf87aedd0ccaa1ae71f7c32d0d0d6e9",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, receipts wave",
      "visibility": "published",
      "path": "/writing/the-evidence-trap/",
      "contentHash": "sha256:e75470768a3bf0a3",
      "releaseHash": "e0b6407ca3a94f5f4cf88ff0aca88bc15cf87aedd0ccaa1ae71f7c32d0d0d6e9"
    }
  ]
}