brick · v1 · 2026-08-03
Below-threshold design
The threshold criteria are not a forecast, they are capabilities someone decides to build
The constructive inverse of the threshold rule, from machine-guarding interlocks to the tutor that withheld answers. The canonical treatment of below-threshold design.
The threshold rule tells you what you are holding: a system with cross-session goals of its own, the capacity to initiate, resistance to interruption, has crossed from tool to actor, and actors get supervised rather than used. That rule is usually read as a classification problem, as though agency were a tide that rises on its own schedule and the builder’s job were to measure the water. This brick is the constructive inverse. Every one of the threshold criteria is a capability, and a capability is a build decision. Nobody discovers that their system initiates actions autonomously; somebody ships it. Where the threshold criteria are capabilities, the discipline is to treat them as capabilities to withhold.
Industrial safety engineering settled the underlying principle long before software existed. A mechanical press can take a hand off in a tenth of a second, and for decades the answer was training, signage, and asking operators to be careful, which worked exactly as well as asking ever works. The answer that actually ended the injuries was the two-hand control: the machine cannot cycle unless both of the operator’s hands are occupied on buttons away from the die. No rule to remember, no compliance to audit, no trust to extend. The dangerous state is unreachable, and everything a guard would have promised is instead simply true. A capability withheld by construction needs no trust and carries no promise, which is the whole mechanism: a system that cannot initiate, cannot persist goals across sessions, and cannot resist interruption does not have to be asked to behave, and the trust it does not need is the risk you do not carry.
The measured version of this argument now exists, and the corpus carries it in the crutch-effect brick. The same model that damaged learning when it would simply answer largely stopped damaging it when the product withheld the answer and scaffolded the struggle instead. Same weights, different shape, opposite outcome. The harm was never a property of what the model could do; it was a property of what the product let it do, which is what below-threshold design claims about agency generally. Advise, hold, speak back: everything below the line, built without the machinery of the fourth level.
The standing objection is that this is modesty, a decision to fall behind. It is the opposite. Below the threshold you get to make demands that are impossible above it, starting with the tool-maker’s oldest one: a tool whose behavior cannot be predicted from its construction is not a mysterious agent, it is a defective tool, and you may say so and fix it. Predictability, auditability, the design stance itself, all of it is purchasing power that exists only on the near side of the line. A vessel is not supposed to sail itself, and the point was never to keep the vessel small. It is that everything a vessel carries arrives because someone could steer it, and the builders who withhold the sail are the ones whose cargo keeps arriving.