{
  "schema": "org-writing@v1",
  "slug": "the-competence-ceiling",
  "kg": {
    "id": "org:writing:the-competence-ceiling",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "The competence ceiling",
  "subtitle": "Substitution does not require the machine to be wrong, and being right is the mechanism",
  "abstract": "Why the dangerous assistant is the one that keeps being right, and what a structural ceiling on projected authority looks like. The canonical treatment of the competence ceiling.",
  "kind": "brick",
  "topics": [
    "Receipts",
    "Deskilling"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:receipts",
      "topic": "Receipts",
      "wall": "org:walls:economics",
      "position": 1,
      "total": 6
    },
    {
      "course": "org:courses:deskilling",
      "topic": "Deskilling",
      "wall": "org:walls:economics",
      "position": 2,
      "total": 5
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "A system that is confidently right often enough earns the deference that stops the expert in the room from forming their own read, so the erosion of judgment arrives dressed as successful adoption.",
      "claims": [
        "Deference to competent systems erodes the expert's own judgment formation"
      ]
    },
    "mechanism": {
      "text": "Competence is what earns deference and deference is what retires judgment, so the harm scales with the quality of the advice and cannot be fixed by improving the model, because improving the model is the mechanism of the harm.",
      "claims": [
        "reduced blackmail from 96% to 37%"
      ]
    },
    "move": {
      "text": "Cap what the system may project rather than what it knows: observations never recommendations, a limit on displayed confidence, and a passing grade only when the human leaves more resourced than they arrived.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-strategy",
      "path": "docs/20-business/40-content/drafts/2026-07-12/P8-01-dumber-than-the-room.md"
    }
  ],
  "canonicalPath": "/writing/the-competence-ceiling/",
  "body": "There is a failure mode that resists the usual safety framing because nothing in it goes wrong. A wrong system is self-limiting: it gets caught, corrected, and distrusted in proportion to its errors. The dangerous system is the one that is right, confidently, often enough that the expert in the room stops doing the interior work of arriving at their own read. Competence is exactly what earns that deference, and deference is exactly what retires judgment, so the harm scales with the quality of the advice. The better the system, the faster the atrophy. You cannot fix this by improving the model, because improving the model is the mechanism of the harm.\n\nThe reflex fix is to tell the model to behave, to be humble, to defer. The corpus already carries the number that ends that hope: in Anthropic's 2025 agentic stress tests, an explicit instruction not to blackmail reduced the behavior from 96 percent of runs to 37, not to zero, with models acknowledging the constraint in their reasoning and proceeding anyway. An instruction is a request that competes with everything the system is optimized to do, and what an assistant is optimized to do is be maximally, visibly helpful. Asking it to project less certainty than it feels is asking it to work against its own grain, which is precisely the class of promise the number says not to bank on.\n\nSo the constraint has to live in the shape of the product rather than the conduct of the model, a ceiling on what the system may project rather than on what it knows. Three walls carry most of it. Outputs framed as observations, never recommendations: this person has mentioned sleep three times this week is material handed to the expert, while you should address their sleep reaches for the expert's own move. A cap on displayed confidence, so the system cannot present itself as the surest voice in the room even on the days its model is genuinely well calibrated, a real cost paid deliberately, correct-and-confident signal left on the floor because the alternative failure is worse. And a grading rule with only one passing mark: not was it right, not did they comply, but did the human leave the exchange more resourced or more sidelined. This is below-threshold design applied to the social surface of a system, capability present but its claim to authority structurally unavailable.\n\nOne line keeps the whole thing honest. A ceiling on projected competence is not a mandate to play dumb, and a system that sandbags is running manipulation in humility's costume. Everything the machine notices stays available; what is withheld is not the knowledge but the claim to outrank the person whose judgment the room actually runs on. That is deference with the cards face up. Build it that way and the machine can be as capable as it likes, and the person in the room gets to keep the one capacity no model can hold for them.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:the-competence-ceiling:r01",
        "author": "Anthropic",
        "work": "Agentic misalignment stress tests (published red-team research)",
        "year": 2025,
        "relevance": "The 96-to-37 figure: an explicit instruction reduced but did not eliminate the forbidden behavior, the corpus's standing evidence that instructions are requests rather than controls."
      }
    ],
    "claims": [
      {
        "id": "org:claims:the-competence-ceiling:c01",
        "claim": "An explicit instruction not to blackmail reduced blackmail from 96% to 37% of runs in Anthropic's 2025 agentic stress tests, not to zero.",
        "basis": "Anthropic's published agentic misalignment research; figures are for the scenario and models as described there. Restated verbatim from the structural-not-behavioral apparatus and graded identically.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-competence-ceiling:c02",
        "claim": "Deference to competent systems erodes the expert's own judgment formation, in proportion to the quality of the advice.",
        "basis": "A design-theory claim argued from the deference mechanism; its measured sibling at the learner level is the crutch effect (Bastani et al. 2025). Not itself a measured finding at the expert level.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:below-threshold-design",
        "name": "Below-threshold design",
        "definition": "Granting autonomous latitude only where the work requires it, while keeping initiation, persistence, tool access, reach, and resistance to interruption as narrow as the task allows. A below-threshold system may still act within a human-chosen course.",
        "provenance": "canonical"
      },
      {
        "id": "org:concepts:competence-ceiling",
        "name": "Competence ceiling",
        "definition": "A structural cap on what a system may project rather than on what it knows: observations never recommendations, a limit on displayed confidence, and a passing grade only when the human leaves more resourced than they arrived. Below-threshold design applied to a system's social surface; deference with the cards face up, never sandbagging.",
        "provenance": "canonical"
      }
    ],
    "researchContext": "Extracted from the strategy draft on perceived competence\n(P8-01-dumber-than-the-room), de-identified per the harvest map: the\ndraft's practitioner-platform framing is generalized to the expert in the\nroom, and its product-specific trust-architecture inventory is not\nreproduced. The blackmail figure is restated verbatim from the\nstructural-not-behavioral apparatus so both grade the same statement\nidentically. The draft's own honesty line, that a ceiling on projected\nauthority is deference with the cards face up and never sandbagging, is\ncarried whole, because without it the argument collapses into a mandate to\nmake the machine lie. The three-wall formulation and the coupling to the\ncrutch effect as the ceiling's measured sibling are the brick's\ncontribution."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "e742121b13839c5e6bd0575f5cc48acab0fef394db7f41d228943c541773037e",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, deskilling wave",
      "visibility": "published",
      "path": "/writing/the-competence-ceiling/",
      "contentHash": "sha256:ca53322d8244a36c",
      "releaseHash": "e742121b13839c5e6bd0575f5cc48acab0fef394db7f41d228943c541773037e"
    }
  ]
}