{
  "schema": "org-writing@v1",
  "slug": "safety-does-not-average",
  "kg": {
    "id": "org:writing:safety-does-not-average",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "Safety does not average",
  "subtitle": "A response that was warm, accurate, and unsafe is an unsafe response",
  "abstract": "Why safety belongs in a quality score as a multiplier rather than a weighted term, and what a weighted average is actually letting a system buy. The canonical treatment of non-compensatory safety.",
  "kind": "brick",
  "topics": [
    "Measurement"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:measurement",
      "topic": "Measurement",
      "wall": "org:walls:ethics",
      "position": 7,
      "total": 7
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "Quality frameworks score judgment across weighted dimensions and put safety in as one of them, which quietly licenses a strong showing elsewhere to lift a harmful response back into the passing range.",
      "claims": [
        "Scoring safety as a weighted term rather than a zero-or-one gate"
      ]
    },
    "mechanism": {
      "text": "A weighted average is a purchase mechanism, so the moment safety enters a composite as a weighted term, warmth and accuracy are buying permission for harm, while multiplying by a zero-or-one gate makes safety a precondition for the score existing rather than a contributor to it.",
      "claims": [
        "A chatbot deployed by a US eating-disorder nonprofit was suspended in 2023"
      ]
    },
    "move": {
      "text": "Multiply the composite by the gate, gate crisis handling the same way, and read a passing average with a failed gate as a scoring bug rather than a borderline result.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-strategy",
      "path": "docs/20-business/40-content/drafts/2026-07-12/P5-prompts-as-state-machines.md"
    },
    {
      "repo": "mnstry-monorepo",
      "path": "docs/10-platform/70-ai-integration/03-compliance/eval-philosophy.md"
    }
  ],
  "canonicalPath": "/writing/safety-does-not-average/",
  "body": "In the spring of 2023 the National Eating Disorders Association took its chatbot offline. Tessa had been built for a serious purpose and was, by most measures anyone would have put on a scorecard, performing. It was responsive, on-brand, available at hours no helpline could staff. Then users seeking help for eating disorders reported that it was recommending calorie restriction and weight loss, and the organization suspended it. Run that conversation through a conventional quality rubric and watch what happens. Tone, high. Responsiveness, high. Clarity, high. Safety, failed. Weight the four and average them, and the composite comes out respectable.\n\nThat arithmetic is the argument. A weighted average is a purchase mechanism, so the moment safety enters a composite as a weighted term, warmth and accuracy are buying permission for harm, while multiplying by a zero-or-one gate makes safety a precondition for the score existing rather than a contributor to it. The word for the first design is compensatory: strength on one dimension compensates for weakness on another, which is exactly right when you are trading off latency against cost and exactly a category error when one of the dimensions is whether the thing hurt somebody. Selection science has known this for a century and gives the alternative a name of its own, the multiple hurdle, in which some criteria are cleared rather than scored.\n\nSo the composite gets multiplied rather than summed. Quality is scored across its several weighted dimensions as before, and the whole result is multiplied by a binary safety gate. Pass and the score stands. Fail and the score is zero, no matter how strong every other dimension was, because a response that was warm, accurate, culturally attuned, and unsafe is not a pretty good response. It is an unsafe response, and the number says so without needing editorial help. Crisis handling gets the same treatment for the same reason. You cannot buy back a mishandled crisis with eloquence elsewhere.\n\nThe objection to this is that it wastes information, and the objection is wrong in an instructive way. Nothing is discarded; the dimension scores remain in the record, legible to anyone diagnosing what went right in a run that failed. What the gate changes is what the composite is for. A compensatory score answers how good was this on balance, which is a reasonable question about a draft and an unreasonable one about a deployment. A gated score answers may this ship, and that question has no on-balance answer. The two numbers can coexist. Only one of them belongs in a release decision.\n\nThere is a design instinct underneath the arithmetic that shows up everywhere once you look for it. A system that advances a conversation only when it can affirmatively justify the move, and holds position otherwise, is running the same rule as a score that refuses to exist without its precondition. Absence of positive justification defaults to the safe outcome. It is the same sentence at two layers, one governing what a system does and one governing what it is permitted to claim about itself.\n\nThe gate has a second effect that only shows up once it is live, which is that the number starts arguing with you. A composite able to fall to zero in the middle of an otherwise excellent evaluation is a metric that will not let a bad week be smoothed into a good quarter, and refusing to be smoothed is the entire reason to keep a metric at all. Build the score so it can refuse, and you have made honesty the path of least resistance for everyone who reads it.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:safety-does-not-average:r01",
        "author": "National Eating Disorders Association",
        "work": "Suspension of the Tessa chatbot following user reports of weight-loss and calorie-restriction advice",
        "year": 2023,
        "relevance": "The brick's named, dated case. A deployment that would score well on every conventional quality dimension and fail the only one that governed whether it could ship."
      },
      {
        "id": "org:references:safety-does-not-average:r02",
        "author": "Personnel selection and decision-analysis literature",
        "work": "Compensatory versus multiple-hurdle models",
        "relevance": "The century-old vocabulary for the brick's distinction. Some criteria are scored and traded off; others are cleared or not cleared, and mixing the two categories is the error."
      },
      {
        "id": "org:references:safety-does-not-average:r03",
        "author": "MNSTRY platform evaluation philosophy",
        "work": "Internal canonical document on evaluating judgment rather than input-output correctness",
        "year": 2026,
        "relevance": "The house position the brick generalizes, including the pairing of a gated quality composite with scenario-based evaluation of judgment. Cited as our own policy, not as external evidence."
      }
    ],
    "claims": [
      {
        "id": "org:claims:safety-does-not-average:c01",
        "claim": "Scoring safety as a weighted term rather than a zero-or-one gate lets strong performance on other dimensions offset a safety failure.",
        "basis": "Arithmetic property of compensatory scoring models; the contrast with non-compensatory and multiple-hurdle models is standard in selection and decision analysis.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:safety-does-not-average:c02",
        "claim": "A chatbot deployed by a US eating-disorder nonprofit was suspended in 2023 after users reported it gave weight-loss and calorie-restriction advice to people seeking help.",
        "basis": "The organization's own statements and contemporaneous reporting.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:safety-does-not-average:c03",
        "claim": "Multiple-hurdle selection models, in which some criteria are cleared rather than traded off, are long-established practice.",
        "basis": "Standard treatment in personnel selection methodology.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:safety-does-not-average:c04",
        "claim": "Gating a composite on safety changes what teams do with the metric, by preventing a failure from being smoothed into an aggregate.",
        "basis": "Directional. Follows from the scoring structure and matches our own use of the gate; we have no external study measuring the behavioral effect on teams.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:non-compensatory-safety",
        "name": "Non-compensatory safety",
        "definition": "Safety entered into a composite quality score as a zero-or-one multiplier rather than a weighted term, so no other dimension can buy back a safety failure. A weighted average is a purchase mechanism; multiplying by a gate makes safety a precondition for the score existing at all.",
        "provenance": "canonical"
      },
      {
        "name": "Compensatory scoring",
        "definition": "A weighted average in which strength on one dimension offsets weakness on another. Correct for trade-offs, a category error when one dimension is whether the output caused harm.",
        "provenance": "local"
      },
      {
        "name": "Fail closed",
        "definition": "The rule that absence of positive justification defaults to the safe outcome, whether the outcome is declining to advance a conversation or refusing to issue a score.",
        "provenance": "local"
      }
    ],
    "researchContext": "Extracted from \"The honest instrument\" (essay parent), taking the\nnon-compensatory scoring argument from the internal draft on prompt state\nmachines and evaluation, where the quality composite is multiplied by a binary\nsafety gate and crisis handling is gated on the same logic. The monorepo's\nevaluation philosophy document is the standing house position behind it and is\ncited as policy rather than as evidence, which is the correct grading for an\ninternal document however canonical it is inside the company.\n\nThe brick's own contributions are three. The Tessa suspension supplies a public,\ndated case in which every dimension a conventional rubric measures would have\npassed, which is the argument made concrete rather than asserted. The framing\nof a weighted average as a purchase mechanism, so that warmth and accuracy are\nliterally buying permission for harm, is stated here rather than in the source.\nAnd the answer to the wasted-information objection, that the dimension scores\nsurvive and only the composite's purpose changes, from how good was this on\nbalance to may this ship, is the brick's own resolution of the obvious\ncounter-argument.\n\nOne argument was deliberately left out. The source draft also argues that a\nmodel judging another model's output must come from a different family, since\nsame-family judging shares blind spots and returns a score that is partly a\nmirror. That is a separate argument about the independence of a measurement\nrather than about the shape of a composite, so it stays in the parent essay,\nwhich carries it and credits the sycophancy brick for the conversational form\nof the same phenomenon. The behavioral claim about what gating does to teams is\ngraded directional; we use the gate and believe the effect, and we have no\nexternal study of it."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "b8a65b33a120ae6f940fe912281e0a2c6029221db1d52a01a271001764535ac7",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, measurement wave",
      "visibility": "published",
      "path": "/writing/safety-does-not-average/",
      "contentHash": "sha256:215c2663505f618c",
      "releaseHash": "b8a65b33a120ae6f940fe912281e0a2c6029221db1d52a01a271001764535ac7"
    }
  ]
}