Measurement · research may 2026 · published 2026-08-03 · v1 · 3 min read
Worst measured at the edges
An instrument is least reliable exactly where its verdicts do the most damage
How a questionnaire can be precise about an average person and nearly blind about an exceptional one, and why the average reliability figure conceals it. The canonical treatment of edge blindness.
Classical test theory hands every person the same error bar, which is a convenient fiction and a consequential one. Item response theory drops the fiction. By modelling each question separately, how hard it is and how sharply it separates someone who has a trait from someone who does not, it can report how much the whole instrument actually knows at every point along the scale. That report is the test information function, and the first time you see one plotted, the shape is a small shock. It is a hill. Information peaks somewhere near the middle of the trait and falls away toward both ends, which means the standard error does the opposite: narrow where the crowd is, wide out where the people are unusual.
Precision depends on where the questions sit, so a scale assembled around the middle of a trait carries its tightest band where nothing is decided and its widest band at the extremes, which is where selection, screening, and alarm all happen. Nobody convenes a review over an average result. The decisions get made about the person at the top of the distribution and the person at the bottom, and those are exactly the two people the questionnaire was least equipped to describe. An instrument can be entirely sound and still be reporting, in effect, that this candidate is somewhere in the top fifth, give or take a fifth.
Consider what that does to a threshold. A program that admits the top five percent needs the instrument to separate the top five percent from the top fifteen, and at that end of the scale it very often cannot, because there were only ever two or three items difficult enough to discriminate up there. The cutoff still runs. The names still sort. What has actually happened is that a coin flip has been laundered through a decimal point, and the people on both sides of the line have been told something about themselves that the evidence does not contain. The same arithmetic runs at the other end, where a screening tool built to describe ordinary distress is asked to identify the person in danger.
Two disciplines follow, and neither requires abandoning the instrument. The first is reporting the error conditionally rather than on average. A single reliability coefficient averages the hill flat and publishes the mean as though it applied to everyone; the professional testing standards ask for the error at the score in question, which is the number that actually governs whether this particular verdict is safe. The second is building toward the edges on purpose. Adaptive testing exists in large part for this reason: rather than asking everyone the same forty questions, most of which tell it nothing, the instrument selects the next question to be the most informative one it can ask given what it already believes, and stops when its uncertainty is low enough. Two people finish at different lengths, not because one is better but because the instrument reached honest confidence about one of them sooner. Graduate admissions testing moved to computerized adaptive form in the 1990s on exactly this logic.
What the hill really tells you is which people the instrument has not yet earned the right to describe, and that is a gift rather than an embarrassment. A blind spot you can plot is a blind spot you can fix, by writing items hard enough and easy enough to see the ends of the range, or by declining the decision until you have. The alternative is not a better measurement. It is the same ignorance without the map. Any instrument that reports where it stops seeing has told you where to build next, and that is the beginning of an instrument worthy of the people at the edges.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 4
-
Frederic Lord and the item response theory literature. Test information functions and conditional standard errors of measurement
The technical core of the brick: information is a function of trait level, peaks where the items are concentrated, and falls away at the extremes, so the standard error is narrowest in the middle and widest at the ends.
Comment on this source -
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing
The requirement that measurement error accompany reported scores, and the specific expectation of conditional standard errors where precision varies across the score range. The averaged coefficient is what the standards are written against.
Comment on this source -
Rizqy Amelia Zein and Hanif Akhtar (2024). Getting started with the graded response model (International Journal of Psychology)
Working treatment of the polytomous model behind per-item discrimination and thresholds, which is where the shape of the information curve comes from on Likert-type instruments.
Comment on this source -
Educational Testing Service (1993). Introduction of the computerized adaptive form of the Graduate Record Examinations
The named case for adaptivity as a response to edge blindness: selecting the next item to be maximally informative for this examinee rather than administering a fixed form built for the middle.
Comment on this source
Claims and confidence 5
- verified
Item response theory reports precision that varies by trait level, so an instrument can be precise in the middle of a scale and nearly uninformative at its extremes.
Test information functions and conditional standard errors; standard result in the IRT literature.
Respond to this claim - verified
Professional testing standards require reported scores to be accompanied by information about measurement error.
Standards for Educational and Psychological Testing (2014), reliability and score-reporting chapters.
Respond to this claim - verified
A single averaged reliability coefficient conceals the variation of precision across the score range.
Arithmetic property of averaging a conditional quantity; stated directly in the psychometrics literature on marginal versus conditional reliability.
Respond to this claim - verified
The Graduate Record Examinations moved to a computerized adaptive form in the 1990s, selecting items by their informativeness for the individual examinee.
Testing-organization documentation of the transition; the specific adaptive granularity has changed since, and the brick claims only the transition and its rationale.
Respond to this claim - directional
Adaptive administration reaches a target precision in fewer items than a fixed form of equivalent precision.
Directional. Consistently reported across the computerized adaptive testing literature, with the size of the reduction dependent on the item bank and the stopping rule.
Respond to this claim