Skip to content

Machine judgment · research july 2026 · published 2026-08-03 · v1 · 3 min read

The unambiguous half

A rule a machine can only half decide is enforced on that half and reported on the rest

How a practice registry sorts its rules into the mechanically enforced and the merely advisory, and why a check that rules past its own evidence costs more than it catches. The canonical treatment of the unambiguous half.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "An external validation of a widely deployed sepsis prediction model raised alerts on 18 percent of all hospitalized patients."

    verified. Wong et al., JAMA Internal Medicine 2021; figures are for the cohort and site described there. Restated verbatim from the apparatus of the discernment essay and graded identically.

Open the complete evidence in the structured publication.

Automated checks are written as though every rule were fully decidable from the artifact, and a check that fires on cases it cannot read trains the people who read it to dismiss it.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "A design-system practice registry sorts entries into mechanic practices enforced by a named validator check and ruling practices that stay advisory until a neutral validator exists, and a mechanized ruling enforces only the unambiguous half of its contract while reporting the judgement-dependent half as a note."

    verified. The registry's own category field and its stated note, read directly from the artifact on 2026-08-03.

Open the complete evidence in the structured publication.

A rule whose breach is decidable from the artifact alone can be mechanized and failed on, and the half requiring judgment can only be reported, so a mechanized ruling enforces the first half and addresses the second to a person.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Split every rule you are about to automate before you automate it, fail the build on the decidable half, and demote the rest to a note that names a human reader.

Every practice in the design registry behind one of our own sites carries a field that decides what a build is allowed to do about it. A practice marked as a mechanic is enforced by a named validator check, and a violation fails the build. A practice marked as a ruling is advisory prose, binding on the people who work there and invisible to the machine, until somebody can state it precisely enough for a neutral check to exist. The field looks like a maturity ladder, prose waiting at the bottom to be promoted into enforcement at the top. It is not. It is a declaration about which kind of rule each practice is, and the interesting ones turn out to be two kinds at once.

Watch what happened to the type ruling when it was finally mechanized. The rule says type is set by role rather than by arbitrary size, and half of it proved decidable from the file alone. A declared size below the readable floor is a breach, and the check fails the build and names the file and the literal that did it. The other half is not decidable at all. A size above the floor but outside the role scale might be a lapse or might be a logotype that has earned its exception, and no scan of the source can tell those apart. So the check does something that looks like a compromise and is actually a boundary. Below the floor it fails. Above it, it emits a note, addressed to a person, saying that this size sits outside the role scale and asking them to confirm it earns the exception.

The mechanism is the split itself, because a rule whose breach is decidable from the artifact alone can be mechanized and failed on while the half that requires reading the situation can only be reported, and a check that cannot distinguish a defensible case from a breach does not produce findings at all, it produces noise, and it spends the build’s credibility to do it. We call the enforceable part the unambiguous half, and the discipline is refusing to pretend the rest is inside it.

The motion ruling was cut the same way, and the residue was left in the open. Toggling an element’s display property to change state cannot animate, so a state change written that way always pops, and the check fails on it without hesitation. The remainder of that ruling, that reveal motion belongs on a unit’s wrapper rather than on the text inside it, stays prose, with a comment in the check explaining precisely why it stayed prose. A class-name scan cannot tell a standalone lead paragraph, which is correct, from inner text inside a card that is also moving, which is the regression. That comment is not an apology for incomplete automation. It is the finding.

What a check costs when it rules past its evidence has been measured, and the number comes from a hospital rather than a build. The sepsis prediction model embedded in one of the most widely deployed hospital record systems in the United States raised alerts on eighteen percent of everyone admitted, and the clinicians did what people always do with an instrument that fires on situations it cannot read. They learned to click through. Alert fatigue is not a weakness in the humans; it is the accurate response to a machine claiming more authority than its evidence supports, and every noisy check in a build teaches the same lesson at a smaller scale, one dismissed warning at a time, until the morning the real one arrives and is dismissed with the rest. A machine that fails only on its unambiguous half gives up coverage it never honestly had and keeps the thing worth more, which is the right to be believed the next time it says no.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 3
  1. MNSTRY (internal) (2026). Marketing-site practice registry and foundation validator suite

    The primary artifact. Every practice is categorised as a mechanic (enforced by a named validator check) or a ruling (advisory until a neutral validator exists), and the registry's own note states that a mechanized ruling enforces only the unambiguous half of its contract and reports the judgement-dependent half as a note. The type-floor and motion checks are the two worked examples in the brick.

    Comment on this source
  2. Andrew Wong, Erkin Otles, John P. Donnelly and colleagues (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients (JAMA Internal Medicine)

    The alert-burden figure, restated from the discernment essay's apparatus and graded identically. Used here for the cost side of the argument rather than the judgment side.

    Comment on this source
  3. Clinical decision support and human factors literature. The alert fatigue literature (override rates for low-specificity interruptive alerts)

    The documented consequence of an instrument that fires on states it cannot read. The brick's transfer of this finding from clinical alerting to build checks is its own argument, not a measured result.

    Comment on this source
Claims and confidence 5
  1. verified

    An external validation of a widely deployed sepsis prediction model raised alerts on 18 percent of all hospitalized patients.

    Wong et al., JAMA Internal Medicine 2021; figures are for the cohort and site described there. Restated verbatim from the apparatus of the discernment essay and graded identically.

    Respond to this claim
  2. verified

    A design-system practice registry sorts entries into mechanic practices enforced by a named validator check and ruling practices that stay advisory until a neutral validator exists, and a mechanized ruling enforces only the unambiguous half of its contract while reporting the judgement-dependent half as a note.

    The registry's own category field and its stated note, read directly from the artifact on 2026-08-03.

    Respond to this claim
  3. verified

    In the mechanized type ruling, a declared size below the type floor fails the build while a size above the floor but outside the role scale emits a note asking a person to confirm the exception.

    The validator source read directly; the note branch is explicit in the check and its comment gives the reason, that a logotype may earn a size outside the role scale.

    Respond to this claim
  4. verified

    The wrapper-versus-inner-text half of the motion ruling was deliberately left as prose because a class-name scan cannot distinguish a standalone lead paragraph from inner text inside a moving card.

    The comment carried in the check itself, which states that a check unable to tell them apart reports noise rather than findings.

    Respond to this claim
  5. verified

    High volumes of low-specificity interruptive alerts produce high override rates among clinicians, a documented failure mode of clinical decision support.

    The clinical alert fatigue literature, consistent across settings; the effect is well established, the specific override rates vary by system and site.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to The unambiguous half

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The unambiguous half

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.