{
  "schema": "org-writing@v1",
  "slug": "the-contaminated-corpus",
  "kg": {
    "id": "org:writing:the-contaminated-corpus",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "The contaminated corpus",
  "subtitle": "Knowledge assembled by machine fails at the seams, and the seams are the one place nothing is looking",
  "abstract": "What a cleanup of our own machine-built research corpus found at the joins, and why the first audit of the damage was wrong in the same way the damage was. The canonical treatment of seam failure in assembled knowledge.",
  "kind": "brick",
  "topics": [
    "Provenance at machine scale"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:provenance",
      "topic": "Provenance at machine scale",
      "wall": "org:walls:engineering",
      "position": 1,
      "total": 5
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "Corpora built by machine at scale carry defects no check can see, because a document made of two unrelated research results is well formed at every level except meaning.",
      "claims": [
        "removed 13,880 lines across 197 files"
      ]
    },
    "mechanism": {
      "text": "A concatenation of two unrelated documents passes every structural check there is, so the only instrument that detects it is a reader who knows what both halves were supposed to be about, and assembly at machine scale is the condition under which no such reader exists.",
      "claims": [
        "joined two unrelated query results into one document"
      ]
    },
    "move": {
      "text": "Validate at the joins rather than at the file, and ask of any automated damage estimate what faculty it is exercising, because a count produced by the same kind of process that produced the damage inherits its blindness.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-research",
      "path": "_CLEANUP_STATUS.md"
    }
  ],
  "canonicalPath": "/writing/the-contaminated-corpus/",
  "body": "Our own research corpus is assembled by machine, and in February we took a spade to it. The cleanup ledger records 13,880 lines removed across 197 files, 84 of them orphaned backup files left behind by an earlier automation that had failed partway through its own repair. The number that matters is smaller and worse. Twenty-three files carried what the record calls cross-topic concatenation, meaning an ingestion pipeline had joined two unrelated query results into one document with nothing between them. Research on human connection ran to its natural close, and then, with no heading, no separator, not so much as a blank line's hesitation, an analysis of an unrelated market began in the same file under the same title.\n\nNothing catches this. The frontmatter is valid. The markdown parses. The headings nest correctly, the links resolve, the file opens cleanly in every tool that was ever going to open it, and a search returns it as one document because that is precisely what it is. A concatenation of two unrelated documents is well formed at every level a machine checks, so the only instrument that detects the defect is a reader who already knows what both halves were supposed to be about, and assembly at machine scale is exactly the condition under which no such reader exists.\n\nThen the second failure, which is the one worth publishing. The first audit of the damage estimated that roughly 37 percent of 451 files might be contaminated, something near 170 of them, and the remediation plan was sized against that figure. The second audit found 23, which the record puts at six to eight percent. The overcount was not carelessness and it was not caution either. Around 350 files in that corpus carry a deliberate two-tier shape, a short summary above a fuller treatment of the same subject, and the first pass could not distinguish two passages worded differently about one topic from two passages about two topics. That is the same faculty the ingestion pipeline had been missing when it made the mess. The instrument that measured the damage was blind in the way the process that caused it was blind, and it got the answer wrong by a factor of five in the direction that made the corpus look worse than it was.\n\nWe keep the ledger, wrong first estimate included, because a corpus that cannot describe its own damage is asking to be trusted rather than checked, and trust is not a property anyone can audit. The disciplines that follow from it are unglamorous and they are all about joins. Validate at the seam rather than at the file, since the file is not where assembled knowledge breaks. Require an ingestion step to say which query produced which block, so a document with two parents says so before a person has to notice. And read every automated damage estimate as a measurement taken by an instrument that may share a blind spot with the thing it is measuring, which is a reason to grade the auditor rather than a reason to stop auditing. Every seam we found is now a place the next pass knows to look, and a structure that can point at its own bad joins gets sounder each time it is doubted.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:the-contaminated-corpus:r01",
        "author": "MNSTRY research operations",
        "work": "Research cleanup, final status (the corpus's own remediation ledger)",
        "year": 2026,
        "relevance": "The whole case. A dated internal record of a machine-assembled corpus being repaired, itemised by phase, including the scope correction in which the first audit's estimate is recorded as wrong and left in the document."
      }
    ],
    "claims": [
      {
        "id": "org:claims:the-contaminated-corpus:c01",
        "claim": "A cleanup pass on our own machine-assembled research corpus removed 13,880 lines across 197 files, 84 of which were orphaned backup files deleted outright rather than truncated.",
        "basis": "The corpus's own cleanup ledger, dated 2026-02-06, whose three-phase table sums to both figures. A self-reported record of our own estate, checked for internal arithmetic and against its phase detail rather than re-derived from the file tree.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-contaminated-corpus:c02",
        "claim": "Twenty-three files carried genuine cross-topic concatenation, in which an ingestion pipeline joined two unrelated query results into one document with no separation between them.",
        "basis": "The same ledger's third phase, which names the three affected file series and the lines removed from each.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-contaminated-corpus:c03",
        "claim": "The first audit estimated that roughly 37 percent of 451 files might be contaminated, and the second audit found 23, which the record puts at six to eight percent.",
        "basis": "The ledger's scope-correction section, which records both rounds and preserves the superseded estimate.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-contaminated-corpus:c04",
        "claim": "The first estimate ran high because the audit could not distinguish different wording on the same topic from genuinely different topics.",
        "basis": "The ledger's own diagnosis of its earlier round. A self-reported cause rather than an independently established one, which is why the brick treats it as the record's account.",
        "confidence": "directional",
        "sources": []
      },
      {
        "id": "org:claims:the-contaminated-corpus:c05",
        "claim": "Roughly 350 files in the same corpus carry a deliberate two-tier structure, a summary above a fuller treatment of the same topic, which the second audit judged legitimate and left alone.",
        "basis": "The ledger's what-remains section, with three named directories confirmed same-topic.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-contaminated-corpus:c06",
        "claim": "The contamination is attributed to batch processing in the ingestion pipeline, which concatenated two query results without document separation.",
        "basis": "The ledger's root-cause section, which marks each of its three attributions as likely rather than established. Graded to match the record's own hedge.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:seam-failure",
        "name": "Seam failure",
        "definition": "A defect in assembled knowledge located at the join between two pieces rather than inside either of them. The join is syntactically perfect, so every structural check passes and only meaning is broken, which makes the defect detectable solely by a reader who already knows what both halves were supposed to be about. Named from our own corpus, where an ingestion pipeline concatenated unrelated query results into single well-formed documents.",
        "provenance": "canonical"
      },
      {
        "name": "Inherited blindness",
        "definition": "The condition of an audit that shares a faculty gap with the process it audits, so its error runs in the same direction as the damage. Named here from the case in which a machine estimate of machine-made contamination overshot by a factor of five.",
        "provenance": "local"
      }
    ],
    "researchContext": "Sourced entirely from our own remediation ledger, which is the point of the\nbrick: the argument is credible because it is an admission, and it would be\nworth nothing sourced from someone else's corpus. Every figure was read\nfrom the file before authoring and checked for internal consistency, phase\nby phase, since a self-reported cleanup is exactly the kind of record that\ndeserves the arithmetic done twice. The two attributions the ledger itself\nhedges, the root cause and the diagnosis of the first audit's overcount,\nare graded directional here to match its hedge rather than hardened into\nfindings by republication.\n\nThe subject corpus is described only as a machine-assembled research\ncorpus. Its domain, contents, and directory structure are not the argument\nand are not disclosed.\n\nThe seam framing, the reading of the overcount as an inherited blindness\nrather than as carelessness, and the move toward validating joins rather\nthan files are the brick's contribution. The signed-stones brick owns\nadmission control at the point of ingestion, which is the discipline that\nwould have prevented this failure upstream; this brick owns what the\nfailure looks like from inside a corpus that already has it."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "ceb702a9934e84fa87f473ef5634de684de7b90dfad8ca2a52b424c6b6686162",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, provenance wave",
      "visibility": "published",
      "path": "/writing/the-contaminated-corpus/",
      "contentHash": "sha256:08e844b3308802df",
      "releaseHash": "ceb702a9934e84fa87f473ef5634de684de7b90dfad8ca2a52b424c6b6686162"
    }
  ]
}