Care · research august 2026 · published 2026-08-03 · v1 · 3 min read
The tautology trap
A scale whose items restate its own outcome measures agreement with itself and reports it as evidence
Why an instrument validated against a paraphrase of itself proves nothing, from a 1950 taxonomy of criterion bias to the evaluation that asks a rater whether an answer was helpful. The canonical treatment of the tautology trap.
Validating an instrument against an outcome is the right instinct, and it is one of the few places a measurement of an interior life can earn its keep. Build the scale, find something in the world it ought to predict, correlate the two, publish the coefficient. The instinct fails in one specific way, and the failure is old enough to carry a proper name and current enough to be shipping inside evaluation harnesses this quarter.
Hubert Brogden and Erwin Taylor set out the taxonomy in Educational and Psychological Measurement in 1950. A criterion, they argued, can be deficient by omitting parts of what it is meant to capture, it can carry scale-unit bias, and it can be contaminated, which is what happens when something extraneous to the intended construct enters the criterion score. The most damaging entrant is the predictor itself. Their prevention rule was correspondingly strict, that nobody assigning criterion ratings may have any knowledge of the test scores, because a supervisor who has seen an aptitude result and then rates performance yields a validity coefficient that is partly a measurement of the test’s influence on the supervisor.
Push that leak to its limit and you arrive at the tautology. When a scale’s items and the criterion it is validated against are drawn from the same construct, the correlation between them was fixed at the moment the items were written, so a validation that looks strong is the instrument meeting its own reflection. No biased rater is required. Nothing has to be contaminated in transit, because the two quantities being compared were never independent to begin with. A scale asking whether you feel calm, correlated against a separate measure of calm, will report a handsome coefficient and will have learned nothing whatever about the world it was supposed to be measuring.
The best documented case in the wild is not a psychological scale at all. Michael Bastedo and Nicholas Bowman modeled the U.S. News college rankings in the American Journal of Education in 2010 and found that a school’s published ranking significantly affected the peer assessment scores it received in subsequent years, independent of changes in institutional quality and even of prior reputation. Peer assessment is an input to the ranking. So part of each year’s ranking is a measurement of the previous year’s ranking, laundered through the impressions of administrators who read it, and the stability everyone cites as evidence of the instrument’s validity is partly evidence of a loop.
Our own field has built that loop and calls it evaluation. A harness that shows a rater a system’s output, asks whether the response was helpful, and reports the aggregate as evidence that the system is helpful has validated nothing. It asked a question and printed the answer back with a different label attached. The criterion and the item are one sentence apart. Add that the rater usually sees only the response, without the context in which unhelpfulness would become visible, and the circle tightens rather than loosening, because the wording being scored has already been optimized against the judgment being solicited.
The discipline is unglamorous and it works. Write the criterion before the items, and require the criterion to be something obtainable by a person who has never seen the instrument. Did the user finish the task without asking again. Did the ticket reopen. Was the practice still going a month later with nobody prompting it. Criteria of that kind can embarrass an instrument, which is precisely their worth, and a team that arranges for its measures to be embarrassable on purpose is the only kind whose good numbers ever meant anything at all.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 3
-
Hubert E. Brogden, Erwin K. Taylor (1950). The Theory and Classification of Criterion Bias (Educational and Psychological Measurement 10(2))
The taxonomy. Criterion deficiency, criterion contamination, and criterion scale unit bias, with contamination defined as extraneous elements entering the criterion score and the predictor named as a principal source.
Comment on this source -
Michael N. Bastedo, Nicholas A. Bowman (2010). U.S. News & World Report College Rankings, Modeling Institutional Effects on Organizational Reputation (American Journal of Education 116(2))
The documented loop. Published rankings significantly affect the peer assessment scores institutions later receive, independent of quality change and of prior reputation, and peer assessment is itself an input to the ranking.
Comment on this source -
Society for Industrial and Organizational Psychology. Criterion theory and development, standard treatments of criterion contamination in personnel selection
The prevention rule the brick quotes in substance, that nobody assigning criterion ratings may know the predictor scores, and the standard account of how shared variance inflates an observed validity coefficient.
Comment on this source
Claims and confidence 3
- verified
Brogden and Taylor classified criterion bias in 1950 into deficiency, contamination, and scale unit bias, defining contamination as the entry of elements extraneous to the intended construct into the criterion score.
The Theory and Classification of Criterion Bias, Educational and Psychological Measurement 10(2), 1950; verified during the spiritual-harvest wave, 2026-08-03.
Respond to this claim - verified
Criterion contamination inflates an observed validity coefficient because predictor and criterion share variance that has nothing to do with the construct, which is why standard practice forbids anyone assigning criterion ratings from knowing the predictor scores.
Standard treatments in the personnel selection and psychometrics literature descending from Brogden and Taylor; the rule and its rationale are textbook rather than contested.
Respond to this claim - verified
Published U.S. News college rankings significantly affected the peer assessment scores institutions received in subsequent years, independent of changes in institutional quality and of prior reputation assessments, and peer assessment is itself an input to the ranking.
Bastedo and Bowman, American Journal of Education 116(2), 2010, structural equation models; verified against the published abstract and secondary accounts during this wave.
Respond to this claim