Skip to content

Measurement · research february to july 2026 · published 2026-08-03 · v1 · 3 min read

Safety does not average

A response that was warm, accurate, and unsafe is an unsafe response

Why safety belongs in a quality score as a multiplier rather than a weighted term, and what a weighted average is actually letting a system buy. The canonical treatment of non-compensatory safety.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Scoring safety as a weighted term rather than a zero-or-one gate lets strong performance on other dimensions offset a safety failure."

    verified. Arithmetic property of compensatory scoring models; the contrast with non-compensatory and multiple-hurdle models is standard in selection and decision analysis.

Open the complete evidence in the structured publication.

Quality frameworks score judgment across weighted dimensions and put safety in as one of them, which quietly licenses a strong showing elsewhere to lift a harmful response back into the passing range.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "A chatbot deployed by a US eating-disorder nonprofit was suspended in 2023 after users reported it gave weight-loss and calorie-restriction advice to people seeking help."

    verified. The organization's own statements and contemporaneous reporting.

Open the complete evidence in the structured publication.

A weighted average is a purchase mechanism, so the moment safety enters a composite as a weighted term, warmth and accuracy are buying permission for harm, while multiplying by a zero-or-one gate makes safety a precondition for the score existing rather than a contributor to it.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Multiply the composite by the gate, gate crisis handling the same way, and read a passing average with a failed gate as a scoring bug rather than a borderline result.

In the spring of 2023 the National Eating Disorders Association took its chatbot offline. Tessa had been built for a serious purpose and was, by most measures anyone would have put on a scorecard, performing. It was responsive, on-brand, available at hours no helpline could staff. Then users seeking help for eating disorders reported that it was recommending calorie restriction and weight loss, and the organization suspended it. Run that conversation through a conventional quality rubric and watch what happens. Tone, high. Responsiveness, high. Clarity, high. Safety, failed. Weight the four and average them, and the composite comes out respectable.

That arithmetic is the argument. A weighted average is a purchase mechanism, so the moment safety enters a composite as a weighted term, warmth and accuracy are buying permission for harm, while multiplying by a zero-or-one gate makes safety a precondition for the score existing rather than a contributor to it. The word for the first design is compensatory: strength on one dimension compensates for weakness on another, which is exactly right when you are trading off latency against cost and exactly a category error when one of the dimensions is whether the thing hurt somebody. Selection science has known this for a century and gives the alternative a name of its own, the multiple hurdle, in which some criteria are cleared rather than scored.

So the composite gets multiplied rather than summed. Quality is scored across its several weighted dimensions as before, and the whole result is multiplied by a binary safety gate. Pass and the score stands. Fail and the score is zero, no matter how strong every other dimension was, because a response that was warm, accurate, culturally attuned, and unsafe is not a pretty good response. It is an unsafe response, and the number says so without needing editorial help. Crisis handling gets the same treatment for the same reason. You cannot buy back a mishandled crisis with eloquence elsewhere.

The objection to this is that it wastes information, and the objection is wrong in an instructive way. Nothing is discarded; the dimension scores remain in the record, legible to anyone diagnosing what went right in a run that failed. What the gate changes is what the composite is for. A compensatory score answers how good was this on balance, which is a reasonable question about a draft and an unreasonable one about a deployment. A gated score answers may this ship, and that question has no on-balance answer. The two numbers can coexist. Only one of them belongs in a release decision.

There is a design instinct underneath the arithmetic that shows up everywhere once you look for it. A system that advances a conversation only when it can affirmatively justify the move, and holds position otherwise, is running the same rule as a score that refuses to exist without its precondition. Absence of positive justification defaults to the safe outcome. It is the same sentence at two layers, one governing what a system does and one governing what it is permitted to claim about itself.

The gate has a second effect that only shows up once it is live, which is that the number starts arguing with you. A composite able to fall to zero in the middle of an otherwise excellent evaluation is a metric that will not let a bad week be smoothed into a good quarter, and refusing to be smoothed is the entire reason to keep a metric at all. Build the score so it can refuse, and you have made honesty the path of least resistance for everyone who reads it.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 3
  1. National Eating Disorders Association (2023). Suspension of the Tessa chatbot following user reports of weight-loss and calorie-restriction advice

    The brick's named, dated case. A deployment that would score well on every conventional quality dimension and fail the only one that governed whether it could ship.

    Comment on this source
  2. Personnel selection and decision-analysis literature. Compensatory versus multiple-hurdle models

    The century-old vocabulary for the brick's distinction. Some criteria are scored and traded off; others are cleared or not cleared, and mixing the two categories is the error.

    Comment on this source
  3. MNSTRY platform evaluation philosophy (2026). Internal canonical document on evaluating judgment rather than input-output correctness

    The house position the brick generalizes, including the pairing of a gated quality composite with scenario-based evaluation of judgment. Cited as our own policy, not as external evidence.

    Comment on this source
Claims and confidence 4
  1. verified

    Scoring safety as a weighted term rather than a zero-or-one gate lets strong performance on other dimensions offset a safety failure.

    Arithmetic property of compensatory scoring models; the contrast with non-compensatory and multiple-hurdle models is standard in selection and decision analysis.

    Respond to this claim
  2. verified

    A chatbot deployed by a US eating-disorder nonprofit was suspended in 2023 after users reported it gave weight-loss and calorie-restriction advice to people seeking help.

    The organization's own statements and contemporaneous reporting.

    Respond to this claim
  3. verified

    Multiple-hurdle selection models, in which some criteria are cleared rather than traded off, are long-established practice.

    Standard treatment in personnel selection methodology.

    Respond to this claim
  4. directional

    Gating a composite on safety changes what teams do with the metric, by preventing a failure from being smoothed into an aggregate.

    Directional. Follows from the scoring structure and matches our own use of the gate; we have no external study measuring the behavioral effect on teams.

    Respond to this claim

Read next

You have walked Being measured end to end: the missing width, the borrowed ruler, the blind edges, the uneven doubt, the two timelines, the reconstructable moment, and the gate that refuses to average. A person is not a score, and the instrument that understands this best is the one that can tell you exactly how much of you its number failed to hold.

Practice Take one score someone has given you, or one your product gives other people, and write three lines about it. What is its missing error bar, honestly estimated. Whose answers were used to build the scale it sits on. And what a person would have to do to appeal it, including whether the record could still show what it said before. If any of the three cannot be answered, you have found the work, and it is smaller than it looks.

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to Safety does not average

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target Safety does not average

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.