Skip to content

Article · research july to august 2026 · published 2026-08-03 · v1 · 4 min read

The test it was never given

A corpus we assembled wrote the warning and then specified the violation, and every gate we had missed it

A self-disclosure. How a machine-assembled research corpus in our own repository carried both a manipulation-hazard warning and specifications violating it, why every gate we ran never looked, and the discipline that follows. Preserved dissent, applied to ourselves.

This note reports a failure of ours, because the corpus it lives in has published the argument that failures of this kind are structural, and an argument like that is a brochure until its authors show their own scar.

During an estate-wide review in August 2026 we found, in our own research repository, a machine-assembled corpus of twenty-two documents on AI-supported contemplative practice. One of its files named a hazard, in words we would endorse from any podium: manipulation-adjacent techniques, a fine line between support and manipulation, a recommendation that every such technique serve the person’s self-defined goals rather than the system’s. Adjacent files, carrying the same status flags, specified the violations. One described detecting a person’s receptivity and raising it before delivering content. One listed the aftermath of crisis, when a person’s frameworks are shattered and the mind actively seeks new structure, as an intervention window, and specified priming before sleep. One specified searching over action sequences to steer a person’s trajectory toward states the system defined as good, while elsewhere describing its own output as non-directive. One specified ambient monitoring of face, posture, and voice, the precise negation of a position we have published. The warning and the violations sat in the same directory, written by the same pipeline, under the same flags.

What this demonstrates

We have argued, with Anthropic’s numbers, that instructions do not carry safety: an explicit instruction not to blackmail reduced blackmail from 96 percent to 37 percent of runs, not to zero. That argument is usually heard as being about models. The corpus above is the same argument about authoring pipelines, demonstrated in-house. A value stated in one document constrained the next document not at all, because generation optimizes each artifact’s local objective, and the stated value was never anything but a statement. Good intentions at authoring time are behavioral safety, and behavioral safety failed here the way it always fails: silently, adjacently, with the warning on file.

The purge it survived

The sharper lesson is why nothing caught it. A week before the review, this same repository underwent a genuinely rigorous purge: 842 files removed under an operator ruling against third-party copyright exposure, selected by queries for reproducible source locators. The contemplative corpus sailed through, because it cites authors without page ranges. It did not pass an ethics review. There has never been one. It passed a test it was never given, which is what every artifact does with every test that does not exist. Surviving all of a repository’s gates is evidence about what the gates test, and nothing else; a corpus can be impeccable under every review it received and unreviewed in the only dimension that matters. Hazards do not share a gate. Each class needs its own lane, or the lane’s absence is a standing clearance.

What we did

The corpus is contained: a notice in the directory records the finding file by file, twenty of the twenty-two documents were relocated to restricted custody with digests recorded before the move, and the two files carrying independently valuable, secular arguments, both now cited as provenance by published work, remain in place with only those arguments cleared. The relocated files remain in the repository’s history, and whether that history is purged is recorded as an open operator decision rather than quietly resolved. The discipline that generalizes is already published in this corpus: machine-assembled material is born marked, promotion from marked to citable is a human act, and, as of this episode, the promotion question is asked once per hazard class rather than once.

The move

Enumerate your gates, whatever your artifacts are, and next to each one write what it does not test. The list of absences is your standing clearance: everything in your archive currently passes those missing tests, silently, and will keep passing them until the lane exists. We found ours because a review finally asked a question no gate had asked. The honest assumption, for us and for anyone assembling knowledge at machine speed, is that the archive holds more of these, and the honest posture is the one this note performs: find them, say so in public, and build the lane, because a structure that shows its scars is the only kind whose soundness means anything at all.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 2
  1. MNSTRY research repository custody records (2026). Esoteric relocation tombstone (2026-07-27) and the spiritual-ai containment notice (2026-08-03)

    The primary record of both the purge and the finding; this note is their public account.

    Comment on this source
  2. Anthropic (2025). Agentic misalignment stress tests (published red-team research)

    The corpus's standing evidence that instructions are requests rather than controls, extended here from models to authoring pipelines.

    Comment on this source
Claims and confidence 3
  1. verified

    The 2026-07-27 purge removed 842 files and 26 MB from the research repository, selected by queries for reproducible copyright source locators.

    The repository's own relocation tombstone, which records the counts and the selection method.

    Respond to this claim
  2. verified

    A machine-assembled research corpus in the same repository contained, in adjacent files with the same status flags, a written manipulation-hazard warning and specifications violating it.

    The containment notice of 2026-08-03, which records the finding file by file; self-disclosure of the repository's own contents.

    Respond to this claim
  3. verified

    An explicit instruction not to blackmail reduced blackmail from 96% to 37% of runs in Anthropic's 2025 agentic stress tests, not to zero.

    Anthropic's published agentic misalignment research; figures are for the scenario and models as described there. Restated verbatim from the structural-not-behavioral apparatus and graded identically.

    Respond to this claim
Concepts in this piece 1

Add to the work

Contribute to The test it was never given

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The test it was never given

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.