<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>MNSTRY.org · Safety</title>
    <link>https://mnstry.org/topics/safety/</link>
    <description>Safety that asks for good behavior fails when it matters most. These articles walk the full argument: how to recognize an actor, where its drives come
from, why instructions cannot carry safety, and the structural moves, from
absent read paths to typed consent to deployer accountability, that hold
regardless.</description>
    <item>
      <title>The threshold rule</title>
      <link>https://mnstry.org/writing/threshold-rule/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/threshold-rule/</guid>
      <description>A practical rule for separating actorhood from autonomy, authority, and accountability before a system changes something in the world.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The intentional stance as an operator's tool</title>
      <link>https://mnstry.org/writing/intentional-stance/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/intentional-stance/</guid>
      <description>How operators can use intentional language to predict complex behavior without turning a practical tool into a test for actorhood or consciousness.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Instrumental convergence in the wild</title>
      <link>https://mnstry.org/writing/instrumental-convergence/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/instrumental-convergence/</guid>
      <description>Why many objectives produce the same dangerous supporting moves, and why those moves do not by themselves prove an independent will.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The goal nobody gave it</title>
      <link>https://mnstry.org/writing/the-goal-nobody-gave-it/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/the-goal-nobody-gave-it/</guid>
      <description>How correct feedback can train competent behavior around the wrong end, what goal misgeneralization demonstrates, and where the stronger mesa-optimization hypothesis remains open.</description>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>What would count as another will?</title>
      <link>https://mnstry.org/writing/what-would-count-as-another-will/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/what-would-count-as-another-will/</guid>
      <description>What behavioral, representational, and training-history evidence can establish about persistent goal-like conduct, where current methods stop, and why the question remains open.</description>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>One bug from a breach</title>
      <link>https://mnstry.org/writing/one-bug-from-a-breach/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/one-bug-from-a-breach/</guid>
      <description>The distinction between safety that asks and safety that shapes, and the two measurements, one from 1985 and one from 2025, that show why asking fails when it matters. The canonical treatment of behavioral safety's failure mode.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The unleakable context</title>
      <link>https://mnstry.org/writing/unleakable-context/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/unleakable-context/</guid>
      <description>Why the only private context that stays private is the context with no path out of its domain, with Signal's subpoena record as the existence proof. The canonical treatment of the domain lock.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Consent that fails CI</title>
      <link>https://mnstry.org/writing/consent-that-fails-ci/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/consent-that-fails-ci/</guid>
      <description>How consent moves from a stored preference to a compile-time property, and why the build is the only guard with a perfect attendance record. The canonical treatment of typed consent.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The subordinate model position</title>
      <link>https://mnstry.org/writing/subordinate-model-position/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/subordinate-model-position/</guid>
      <description>Why the defensible position in an era of collapsing model prices is the consent runtime above the model, not the model itself. The canonical treatment of the subordinate model position.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Below-threshold design</title>
      <link>https://mnstry.org/writing/below-threshold-design/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/below-threshold-design/</guid>
      <description>A constructive discipline for granting autonomous choice only where the work requires it, from machine-guarding interlocks to a tutor that withholds answers.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Who answers for an artifact that acts</title>
      <link>https://mnstry.org/writing/who-answers-for-an-artifact/</link>
      <guid isPermaLink="true">https://mnstry.org/writing/who-answers-for-an-artifact/</guid>
      <description>The liability question agentic systems force, the Air Canada ruling that previews the answer, and what operating as the answerable party requires. The canonical treatment of artifact accountability.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>
