Skip to content

Provenance at machine scale · research may 2026 · published 2026-08-03 · v1 · 3 min read

The contaminated corpus

Knowledge assembled by machine fails at the seams, and the seams are the one place nothing is looking

What a cleanup of our own machine-built research corpus found at the joins, and why the first audit of the damage was wrong in the same way the damage was. The canonical treatment of seam failure in assembled knowledge.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "A cleanup pass on our own machine-assembled research corpus removed 13,880 lines across 197 files, 84 of which were orphaned backup files deleted outright rather than truncated."

    verified. The corpus's own cleanup ledger, dated 2026-02-06, whose three-phase table sums to both figures. A self-reported record of our own estate, checked for internal arithmetic and against its phase detail rather than re-derived from the file tree.

Open the complete evidence in the structured publication.

Corpora built by machine at scale carry defects no check can see, because a document made of two unrelated research results is well formed at every level except meaning.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Twenty-three files carried genuine cross-topic concatenation, in which an ingestion pipeline joined two unrelated query results into one document with no separation between them."

    verified. The same ledger's third phase, which names the three affected file series and the lines removed from each.

Open the complete evidence in the structured publication.

A concatenation of two unrelated documents passes every structural check there is, so the only instrument that detects it is a reader who knows what both halves were supposed to be about, and assembly at machine scale is the condition under which no such reader exists.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Validate at the joins rather than at the file, and ask of any automated damage estimate what faculty it is exercising, because a count produced by the same kind of process that produced the damage inherits its blindness.

Our own research corpus is assembled by machine, and in February we took a spade to it. The cleanup ledger records 13,880 lines removed across 197 files, 84 of them orphaned backup files left behind by an earlier automation that had failed partway through its own repair. The number that matters is smaller and worse. Twenty-three files carried what the record calls cross-topic concatenation, meaning an ingestion pipeline had joined two unrelated query results into one document with nothing between them. Research on human connection ran to its natural close, and then, with no heading, no separator, not so much as a blank line’s hesitation, an analysis of an unrelated market began in the same file under the same title.

Nothing catches this. The frontmatter is valid. The markdown parses. The headings nest correctly, the links resolve, the file opens cleanly in every tool that was ever going to open it, and a search returns it as one document because that is precisely what it is. A concatenation of two unrelated documents is well formed at every level a machine checks, so the only instrument that detects the defect is a reader who already knows what both halves were supposed to be about, and assembly at machine scale is exactly the condition under which no such reader exists.

Then the second failure, which is the one worth publishing. The first audit of the damage estimated that roughly 37 percent of 451 files might be contaminated, something near 170 of them, and the remediation plan was sized against that figure. The second audit found 23, which the record puts at six to eight percent. The overcount was not carelessness and it was not caution either. Around 350 files in that corpus carry a deliberate two-tier shape, a short summary above a fuller treatment of the same subject, and the first pass could not distinguish two passages worded differently about one topic from two passages about two topics. That is the same faculty the ingestion pipeline had been missing when it made the mess. The instrument that measured the damage was blind in the way the process that caused it was blind, and it got the answer wrong by a factor of five in the direction that made the corpus look worse than it was.

We keep the ledger, wrong first estimate included, because a corpus that cannot describe its own damage is asking to be trusted rather than checked, and trust is not a property anyone can audit. The disciplines that follow from it are unglamorous and they are all about joins. Validate at the seam rather than at the file, since the file is not where assembled knowledge breaks. Require an ingestion step to say which query produced which block, so a document with two parents says so before a person has to notice. And read every automated damage estimate as a measurement taken by an instrument that may share a blind spot with the thing it is measuring, which is a reason to grade the auditor rather than a reason to stop auditing. Every seam we found is now a place the next pass knows to look, and a structure that can point at its own bad joins gets sounder each time it is doubted.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 1
  1. MNSTRY research operations (2026). Research cleanup, final status (the corpus's own remediation ledger)

    The whole case. A dated internal record of a machine-assembled corpus being repaired, itemised by phase, including the scope correction in which the first audit's estimate is recorded as wrong and left in the document.

    Comment on this source
Claims and confidence 6
  1. verified

    A cleanup pass on our own machine-assembled research corpus removed 13,880 lines across 197 files, 84 of which were orphaned backup files deleted outright rather than truncated.

    The corpus's own cleanup ledger, dated 2026-02-06, whose three-phase table sums to both figures. A self-reported record of our own estate, checked for internal arithmetic and against its phase detail rather than re-derived from the file tree.

    Respond to this claim
  2. verified

    Twenty-three files carried genuine cross-topic concatenation, in which an ingestion pipeline joined two unrelated query results into one document with no separation between them.

    The same ledger's third phase, which names the three affected file series and the lines removed from each.

    Respond to this claim
  3. verified

    The first audit estimated that roughly 37 percent of 451 files might be contaminated, and the second audit found 23, which the record puts at six to eight percent.

    The ledger's scope-correction section, which records both rounds and preserves the superseded estimate.

    Respond to this claim
  4. directional

    The first estimate ran high because the audit could not distinguish different wording on the same topic from genuinely different topics.

    The ledger's own diagnosis of its earlier round. A self-reported cause rather than an independently established one, which is why the brick treats it as the record's account.

    Respond to this claim
  5. verified

    Roughly 350 files in the same corpus carry a deliberate two-tier structure, a summary above a fuller treatment of the same topic, which the second audit judged legitimate and left alone.

    The ledger's what-remains section, with three named directories confirmed same-topic.

    Respond to this claim
  6. directional

    The contamination is attributed to batch processing in the ingestion pipeline, which concatenated two query results without document separation.

    The ledger's root-cause section, which marks each of its three attributions as likely rather than established. Graded to match the record's own hedge.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to The contaminated corpus

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The contaminated corpus

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.