Skip to content

Care · research august 2026 · published 2026-08-03 · v1 · 3 min read

The tautology trap

A scale whose items restate its own outcome measures agreement with itself and reports it as evidence

Why an instrument validated against a paraphrase of itself proves nothing, from a 1950 taxonomy of criterion bias to the evaluation that asks a rater whether an answer was helpful. The canonical treatment of the tautology trap.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Published U.S. News college rankings significantly affected the peer assessment scores institutions received in subsequent years, independent of changes in institutional quality and of prior reputation assessments, and peer assessment is itself an input to the ranking."

    verified. Bastedo and Bowman, American Journal of Education 116(2), 2010, structural equation models; verified against the published abstract and secondary accounts during this wave.

Open the complete evidence in the structured publication.

Validating an instrument against an outcome is the right instinct, and it fails in one specific way that is old enough to have a proper name and current enough to be shipping in evaluation harnesses.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Criterion contamination inflates an observed validity coefficient because predictor and criterion share variance that has nothing to do with the construct, which is why standard practice forbids anyone assigning criterion ratings from knowing the predictor scores."

    verified. Standard treatments in the personnel selection and psychometrics literature descending from Brogden and Taylor; the rule and its rationale are textbook rather than contested.

Open the complete evidence in the structured publication.

When the items and the criterion are drawn from the same construct the correlation between them was fixed at the moment the items were written, so a validation that looks strong is the instrument meeting its own reflection.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Write the criterion before the items and require it to be obtainable by a person who has never seen the instrument.

Validating an instrument against an outcome is the right instinct, and it is one of the few places a measurement of an interior life can earn its keep. Build the scale, find something in the world it ought to predict, correlate the two, publish the coefficient. The instinct fails in one specific way, and the failure is old enough to carry a proper name and current enough to be shipping inside evaluation harnesses this quarter.

Hubert Brogden and Erwin Taylor set out the taxonomy in Educational and Psychological Measurement in 1950. A criterion, they argued, can be deficient by omitting parts of what it is meant to capture, it can carry scale-unit bias, and it can be contaminated, which is what happens when something extraneous to the intended construct enters the criterion score. The most damaging entrant is the predictor itself. Their prevention rule was correspondingly strict, that nobody assigning criterion ratings may have any knowledge of the test scores, because a supervisor who has seen an aptitude result and then rates performance yields a validity coefficient that is partly a measurement of the test’s influence on the supervisor.

Push that leak to its limit and you arrive at the tautology. When a scale’s items and the criterion it is validated against are drawn from the same construct, the correlation between them was fixed at the moment the items were written, so a validation that looks strong is the instrument meeting its own reflection. No biased rater is required. Nothing has to be contaminated in transit, because the two quantities being compared were never independent to begin with. A scale asking whether you feel calm, correlated against a separate measure of calm, will report a handsome coefficient and will have learned nothing whatever about the world it was supposed to be measuring.

The best documented case in the wild is not a psychological scale at all. Michael Bastedo and Nicholas Bowman modeled the U.S. News college rankings in the American Journal of Education in 2010 and found that a school’s published ranking significantly affected the peer assessment scores it received in subsequent years, independent of changes in institutional quality and even of prior reputation. Peer assessment is an input to the ranking. So part of each year’s ranking is a measurement of the previous year’s ranking, laundered through the impressions of administrators who read it, and the stability everyone cites as evidence of the instrument’s validity is partly evidence of a loop.

Our own field has built that loop and calls it evaluation. A harness that shows a rater a system’s output, asks whether the response was helpful, and reports the aggregate as evidence that the system is helpful has validated nothing. It asked a question and printed the answer back with a different label attached. The criterion and the item are one sentence apart. Add that the rater usually sees only the response, without the context in which unhelpfulness would become visible, and the circle tightens rather than loosening, because the wording being scored has already been optimized against the judgment being solicited.

The discipline is unglamorous and it works. Write the criterion before the items, and require the criterion to be something obtainable by a person who has never seen the instrument. Did the user finish the task without asking again. Did the ticket reopen. Was the practice still going a month later with nobody prompting it. Criteria of that kind can embarrass an instrument, which is precisely their worth, and a team that arranges for its measures to be embarrassable on purpose is the only kind whose good numbers ever meant anything at all.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 3
  1. Hubert E. Brogden, Erwin K. Taylor (1950). The Theory and Classification of Criterion Bias (Educational and Psychological Measurement 10(2))

    The taxonomy. Criterion deficiency, criterion contamination, and criterion scale unit bias, with contamination defined as extraneous elements entering the criterion score and the predictor named as a principal source.

    Comment on this source
  2. Michael N. Bastedo, Nicholas A. Bowman (2010). U.S. News & World Report College Rankings, Modeling Institutional Effects on Organizational Reputation (American Journal of Education 116(2))

    The documented loop. Published rankings significantly affect the peer assessment scores institutions later receive, independent of quality change and of prior reputation, and peer assessment is itself an input to the ranking.

    Comment on this source
  3. Society for Industrial and Organizational Psychology. Criterion theory and development, standard treatments of criterion contamination in personnel selection

    The prevention rule the brick quotes in substance, that nobody assigning criterion ratings may know the predictor scores, and the standard account of how shared variance inflates an observed validity coefficient.

    Comment on this source
Claims and confidence 3
  1. verified

    Brogden and Taylor classified criterion bias in 1950 into deficiency, contamination, and scale unit bias, defining contamination as the entry of elements extraneous to the intended construct into the criterion score.

    The Theory and Classification of Criterion Bias, Educational and Psychological Measurement 10(2), 1950; verified during the spiritual-harvest wave, 2026-08-03.

    Respond to this claim
  2. verified

    Criterion contamination inflates an observed validity coefficient because predictor and criterion share variance that has nothing to do with the construct, which is why standard practice forbids anyone assigning criterion ratings from knowing the predictor scores.

    Standard treatments in the personnel selection and psychometrics literature descending from Brogden and Taylor; the rule and its rationale are textbook rather than contested.

    Respond to this claim
  3. verified

    Published U.S. News college rankings significantly affected the peer assessment scores institutions received in subsequent years, independent of changes in institutional quality and of prior reputation assessments, and peer assessment is itself an input to the ranking.

    Bastedo and Bowman, American Journal of Education 116(2), 2010, structural equation models; verified against the published abstract and secondary accounts during this wave.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to The tautology trap

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The tautology trap

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.