{
  "schema": "org-writing@v1",
  "slug": "the-tautology-trap",
  "kg": {
    "id": "org:writing:the-tautology-trap",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "The tautology trap",
  "subtitle": "A scale whose items restate its own outcome measures agreement with itself and reports it as evidence",
  "abstract": "Why an instrument validated against a paraphrase of itself proves nothing, from a 1950 taxonomy of criterion bias to the evaluation that asks a rater whether an answer was helpful. The canonical treatment of the tautology trap.",
  "kind": "brick",
  "topics": [
    "Care"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:instruments-of-care",
      "topic": "Care",
      "wall": "org:walls:ethics",
      "position": 2,
      "total": 5
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "Validating an instrument against an outcome is the right instinct, and it fails in one specific way that is old enough to have a proper name and current enough to be shipping in evaluation harnesses.",
      "claims": [
        "significantly affected the peer assessment scores institutions received in subsequent years"
      ]
    },
    "mechanism": {
      "text": "When the items and the criterion are drawn from the same construct the correlation between them was fixed at the moment the items were written, so a validation that looks strong is the instrument meeting its own reflection.",
      "claims": [
        "share variance that has nothing to do with the construct"
      ]
    },
    "move": {
      "text": "Write the criterion before the items and require it to be obtainable by a person who has never seen the instrument.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-org",
      "path": "editorial/notes/spiritual-corpus-review.md"
    }
  ],
  "canonicalPath": "/writing/the-tautology-trap/",
  "body": "Validating an instrument against an outcome is the right instinct, and it is one of the few places a measurement of an interior life can earn its keep. Build the scale, find something in the world it ought to predict, correlate the two, publish the coefficient. The instinct fails in one specific way, and the failure is old enough to carry a proper name and current enough to be shipping inside evaluation harnesses this quarter.\n\nHubert Brogden and Erwin Taylor set out the taxonomy in Educational and Psychological Measurement in 1950. A criterion, they argued, can be deficient by omitting parts of what it is meant to capture, it can carry scale-unit bias, and it can be contaminated, which is what happens when something extraneous to the intended construct enters the criterion score. The most damaging entrant is the predictor itself. Their prevention rule was correspondingly strict, that nobody assigning criterion ratings may have any knowledge of the test scores, because a supervisor who has seen an aptitude result and then rates performance yields a validity coefficient that is partly a measurement of the test's influence on the supervisor.\n\nPush that leak to its limit and you arrive at the tautology. When a scale's items and the criterion it is validated against are drawn from the same construct, the correlation between them was fixed at the moment the items were written, so a validation that looks strong is the instrument meeting its own reflection. No biased rater is required. Nothing has to be contaminated in transit, because the two quantities being compared were never independent to begin with. A scale asking whether you feel calm, correlated against a separate measure of calm, will report a handsome coefficient and will have learned nothing whatever about the world it was supposed to be measuring.\n\nThe best documented case in the wild is not a psychological scale at all. Michael Bastedo and Nicholas Bowman modeled the U.S. News college rankings in the American Journal of Education in 2010 and found that a school's published ranking significantly affected the peer assessment scores it received in subsequent years, independent of changes in institutional quality and even of prior reputation. Peer assessment is an input to the ranking. So part of each year's ranking is a measurement of the previous year's ranking, laundered through the impressions of administrators who read it, and the stability everyone cites as evidence of the instrument's validity is partly evidence of a loop.\n\nOur own field has built that loop and calls it evaluation. A harness that shows a rater a system's output, asks whether the response was helpful, and reports the aggregate as evidence that the system is helpful has validated nothing. It asked a question and printed the answer back with a different label attached. The criterion and the item are one sentence apart. Add that the rater usually sees only the response, without the context in which unhelpfulness would become visible, and the circle tightens rather than loosening, because the wording being scored has already been optimized against the judgment being solicited.\n\nThe discipline is unglamorous and it works. Write the criterion before the items, and require the criterion to be something obtainable by a person who has never seen the instrument. Did the user finish the task without asking again. Did the ticket reopen. Was the practice still going a month later with nobody prompting it. Criteria of that kind can embarrass an instrument, which is precisely their worth, and a team that arranges for its measures to be embarrassable on purpose is the only kind whose good numbers ever meant anything at all.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:the-tautology-trap:r01",
        "author": "Hubert E. Brogden, Erwin K. Taylor",
        "work": "The Theory and Classification of Criterion Bias (Educational and Psychological Measurement 10(2))",
        "year": 1950,
        "relevance": "The taxonomy. Criterion deficiency, criterion contamination, and criterion scale unit bias, with contamination defined as extraneous elements entering the criterion score and the predictor named as a principal source."
      },
      {
        "id": "org:references:the-tautology-trap:r02",
        "author": "Michael N. Bastedo, Nicholas A. Bowman",
        "work": "U.S. News & World Report College Rankings, Modeling Institutional Effects on Organizational Reputation (American Journal of Education 116(2))",
        "year": 2010,
        "relevance": "The documented loop. Published rankings significantly affect the peer assessment scores institutions later receive, independent of quality change and of prior reputation, and peer assessment is itself an input to the ranking."
      },
      {
        "id": "org:references:the-tautology-trap:r03",
        "author": "Society for Industrial and Organizational Psychology",
        "work": "Criterion theory and development, standard treatments of criterion contamination in personnel selection",
        "relevance": "The prevention rule the brick quotes in substance, that nobody assigning criterion ratings may know the predictor scores, and the standard account of how shared variance inflates an observed validity coefficient."
      }
    ],
    "claims": [
      {
        "id": "org:claims:the-tautology-trap:c01",
        "claim": "Brogden and Taylor classified criterion bias in 1950 into deficiency, contamination, and scale unit bias, defining contamination as the entry of elements extraneous to the intended construct into the criterion score.",
        "basis": "The Theory and Classification of Criterion Bias, Educational and Psychological Measurement 10(2), 1950; verified during the spiritual-harvest wave, 2026-08-03.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-tautology-trap:c02",
        "claim": "Criterion contamination inflates an observed validity coefficient because predictor and criterion share variance that has nothing to do with the construct, which is why standard practice forbids anyone assigning criterion ratings from knowing the predictor scores.",
        "basis": "Standard treatments in the personnel selection and psychometrics literature descending from Brogden and Taylor; the rule and its rationale are textbook rather than contested.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:the-tautology-trap:c03",
        "claim": "Published U.S. News college rankings significantly affected the peer assessment scores institutions received in subsequent years, independent of changes in institutional quality and of prior reputation assessments, and peer assessment is itself an input to the ranking.",
        "basis": "Bastedo and Bowman, American Journal of Education 116(2), 2010, structural equation models; verified against the published abstract and secondary accounts during this wave.",
        "confidence": "verified",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:tautology-trap",
        "name": "Tautology trap",
        "definition": "A validation in which an instrument's items and the criterion it is measured against are drawn from one construct, so the correlation between them was determined at authoring time and reports the wording rather than the world. The limiting case of criterion contamination, which requires a leak between predictor and criterion; the tautology requires no leak, because the two were never independent. Distinct from the borrowed ruler, which concerns whose population calibrated the scale. The contemporary instance is an evaluation that asks a rater whether a response was helpful and reports the aggregate as evidence of helpfulness.",
        "provenance": "canonical"
      }
    ],
    "researchContext": "The review's verification gate on this candidate was mandatory and it is\ndischarged by replacement rather than by confirmation. The source research\nfile states the critique in one clause about spirituality scales whose items\ndirectly measure well-being, and the review's ruling was explicit that the\nfile must not be cited in the published piece, because its remaining\nquantitative content is unusable and a citation invites a reader into it.\nAccordingly the file appears nowhere, including in the brick's source list,\nand the entire argument is rebuilt from methodology literature located during\nthis wave. Not one number from that file is used, and no spiritual-state\nquantification of any kind survives into this brick.\n\nThe Brogden and Taylor taxonomy supplies the lineage and the U.S. News case\nsupplies the anchor, and the case was chosen deliberately from outside\npsychology so that the argument would not read as a quarrel about\npsychometrics. The circularity in the rankings is documented in a\npeer-reviewed model rather than asserted, which is what the argument needed.\n\nThe contemporary application is ours and appears in no source. An evaluation\nthat asks a rater whether an answer was helpful and reports the aggregate as\nevidence of helpfulness is the same circle with the interval between item and\ncriterion reduced to a single sentence, and the observation that the response\nbeing rated has already been optimized against the judgment being solicited\ntightens it further. The same-family judge problem is adjacent and is\ndeliberately not restated, since the parent essay already carries it; this\nbrick is about the identity between item and criterion, which holds even when\nthe rater is human and disinterested."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "5720f7c6c0eb01212fad3670c9ace518aec7e980602722472daf578e4ef3791e",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, spiritual-harvest wave",
      "visibility": "published",
      "path": "/writing/the-tautology-trap/",
      "contentHash": "sha256:fbc0ebd26ca091b8",
      "releaseHash": "5720f7c6c0eb01212fad3670c9ace518aec7e980602722472daf578e4ef3791e"
    }
  ]
}