Topic · 11 articles
Safety
Safety that asks for good behavior fails when it matters most. These articles walk the full argument: how to recognize an actor, where its drives come from, why instructions cannot carry safety, and the structural moves, from absent read paths to typed consent to deployer accountability, that hold regardless.
Start reading The five-minute version
- 1 The threshold rule read
- 2 The intentional stance as an operator's tool read
- 3 Instrumental convergence in the wild read
- 4 The goal nobody gave it read
- 5 What would count as another will? read
- 6 One bug from a breach read
- 7 The unleakable context read
- 8 Consent that fails CI read
- 9 The subordinate model position read
- 10 Below-threshold design read
- 11 Who answers for an artifact that acts read
Practice Pick one system you operate or depend on and place it on the agency spectrum in writing: which of the four threshold criteria does it meet today, and which could it meet under pressure? Then find one rule in your life that works by asking for good behavior, and sketch what removing the path would look like instead.