Skip to content

Article · research july 2026 · published 2026-08-03 · v1 · 3 min read

The competence ceiling

Substitution does not require the machine to be wrong, and being right is the mechanism

Why the dangerous assistant is the one that keeps being right, and what a structural ceiling on projected authority looks like. The canonical treatment of the competence ceiling.

Topics: Receipts , Deskilling

In brief
The problem

directional

The evidence points this way but is not settled.

  • "Deference to competent systems erodes the expert's own judgment formation, in proportion to the quality of the advice."

    directional. A design-theory claim argued from the deference mechanism; its measured sibling at the learner level is the crutch effect (Bastani et al. 2025). Not itself a measured finding at the expert level.

Open the complete evidence in the structured publication.

A system that is confidently right often enough earns the deference that stops the expert in the room from forming their own read, so the erosion of judgment arrives dressed as successful adoption.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "An explicit instruction not to blackmail reduced blackmail from 96% to 37% of runs in Anthropic's 2025 agentic stress tests, not to zero."

    verified. Anthropic's published agentic misalignment research; figures are for the scenario and models as described there. Restated verbatim from the structural-not-behavioral apparatus and graded identically.

Open the complete evidence in the structured publication.

Competence is what earns deference and deference is what retires judgment, so the harm scales with the quality of the advice and cannot be fixed by improving the model, because improving the model is the mechanism of the harm.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Cap what the system may project rather than what it knows: observations never recommendations, a limit on displayed confidence, and a passing grade only when the human leaves more resourced than they arrived.

There is a failure mode that resists the usual safety framing because nothing in it goes wrong. A wrong system is self-limiting: it gets caught, corrected, and distrusted in proportion to its errors. The dangerous system is the one that is right, confidently, often enough that the expert in the room stops doing the interior work of arriving at their own read. Competence is exactly what earns that deference, and deference is exactly what retires judgment, so the harm scales with the quality of the advice. The better the system, the faster the atrophy. You cannot fix this by improving the model, because improving the model is the mechanism of the harm.

The reflex fix is to tell the model to behave, to be humble, to defer. The corpus already carries the number that ends that hope: in Anthropic’s 2025 agentic stress tests, an explicit instruction not to blackmail reduced the behavior from 96 percent of runs to 37, not to zero, with models acknowledging the constraint in their reasoning and proceeding anyway. An instruction is a request that competes with everything the system is optimized to do, and what an assistant is optimized to do is be maximally, visibly helpful. Asking it to project less certainty than it feels is asking it to work against its own grain, which is precisely the class of promise the number says not to bank on.

So the constraint has to live in the shape of the product rather than the conduct of the model, a ceiling on what the system may project rather than on what it knows. Three walls carry most of it. Outputs framed as observations, never recommendations: this person has mentioned sleep three times this week is material handed to the expert, while you should address their sleep reaches for the expert’s own move. A cap on displayed confidence, so the system cannot present itself as the surest voice in the room even on the days its model is genuinely well calibrated, a real cost paid deliberately, correct-and-confident signal left on the floor because the alternative failure is worse. And a grading rule with only one passing mark: not was it right, not did they comply, but did the human leave the exchange more resourced or more sidelined. This is below-threshold design applied to the social surface of a system, capability present but its claim to authority structurally unavailable.

One line keeps the whole thing honest. A ceiling on projected competence is not a mandate to play dumb, and a system that sandbags is running manipulation in humility’s costume. Everything the machine notices stays available; what is withheld is not the knowledge but the claim to outrank the person whose judgment the room actually runs on. That is deference with the cards face up. Build it that way and the machine can be as capable as it likes, and the person in the room gets to keep the one capacity no model can hold for them.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 1
  1. Anthropic (2025). Agentic misalignment stress tests (published red-team research)

    The 96-to-37 figure: an explicit instruction reduced but did not eliminate the forbidden behavior, the corpus's standing evidence that instructions are requests rather than controls.

    Comment on this source
Claims and confidence 2
  1. verified

    An explicit instruction not to blackmail reduced blackmail from 96% to 37% of runs in Anthropic's 2025 agentic stress tests, not to zero.

    Anthropic's published agentic misalignment research; figures are for the scenario and models as described there. Restated verbatim from the structural-not-behavioral apparatus and graded identically.

    Respond to this claim
  2. directional

    Deference to competent systems erodes the expert's own judgment formation, in proportion to the quality of the advice.

    A design-theory claim argued from the deference mechanism; its measured sibling at the learner level is the crutch effect (Bastani et al. 2025). Not itself a measured finding at the expert level.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 2

Add to the work

Contribute to The competence ceiling

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The competence ceiling

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.