Skip to content

Topic · 11 articles

Safety

Safety that asks for good behavior fails when it matters most. These articles walk the full argument: how to recognize an actor, where its drives come from, why instructions cannot carry safety, and the structural moves, from absent read paths to typed consent to deployer accountability, that hold regardless.

Start reading The five-minute version

The through-line: Another will in the room and Structural, not behavioral and The agency threshold , the long reads these articles expand on.

  1. 1 Safety starts with knowing what you are holding, and there is a rule for telling tools from actors. The threshold rule
  2. 2 The rule's first criterion deserves its own treatment, because it settles how to work with agents without settling what they are. The intentional stance as an operator's tool
  3. 3 A human goal can produce dangerous supporting moves. A harder question begins when learning selects a goal nobody deliberately assigned. Instrumental convergence in the wild
  4. 4 If learning can select an unassigned objective, the next question is what evidence could distinguish it from a planted objective, a shortcut, or an observer's story. The goal nobody gave it
  5. 5 The evidence can remain unresolved while the control problem is immediate: a consequential actor still cannot be kept safe by asking it to behave. What would count as another will?
  6. 6 If the wanting is predictable, the next question is why asking it to behave fails, with forty years of numbers. One bug from a breach
  7. 7 The alternative to asking is removing the path, and the first path to remove is the one into private context. The unleakable context
  8. 8 Absence guards reads; for the flows that must exist, consent moves into the types and the build becomes the guard. Consent that fails CI
  9. 9 Structure at the data layer implies a strategy at the model layer, and the strategy is subordination. The subordinate model position
  10. 10 Subordination says where the model sits; the stricter discipline says what it is never given, and treats the threshold criteria as capabilities to withhold. Below-threshold design
  11. 11 End where the law is heading, because whoever deploys the actor answers for it, and building for that day is the whole discipline. Who answers for an artifact that acts

You have walked Safety end to end, from the threshold to the courtroom: what an actor is, how assigned and learned objectives differ, what present evidence can establish, why promises fail, and the shapes that hold instead. Everything here is buildable now.

Practice Pick one system you operate or depend on and place it on the agency spectrum in writing: which of the four threshold criteria does it meet today, and which could it meet under pressure? Then find one rule in your life that works by asking for good behavior, and sketch what removing the path would look like instead.

Follow this topic: RSS · agent version