Skip to content

Measurement · research february to may 2026 · published 2026-08-03 · v1 · 3 min read

Scores without confidence

A judgment that does not carry the width of its own doubt is malformed, not modest

Why a number about a person without its error band is a construction error rather than a rounding convenience, and what the band is actually telling you. The canonical treatment of the malformed score.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Cronbach's alpha is a lower bound that rises with the number of items and does not establish that a scale measures one construct."

    verified. Sijtsma 2009 and the subsequent reliability literature; uncontroversial in psychometrics.

Open the complete evidence in the structured publication.

Instruments routinely publish a bare number about a person, and the most familiar certificate of quality behind such numbers, a high alpha, establishes almost nothing about whether the scale measures one thing.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Professional testing standards require reported scores to be accompanied by information about measurement error."

    verified. Standards for Educational and Psychological Testing (2014), reliability and score-reporting chapters.

Open the complete evidence in the structured publication.

An interval is not commentary on a score but part of it, so a point estimate released without one is not a cautious claim made quietly, it is an incomplete claim made confidently, and everyone downstream treats the missing width as zero.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Treat a score arriving without its error as malformed rather than merely undocumented, and refuse to act on it until the width is supplied.

In August 2010 the Los Angeles Times published effectiveness ratings for thousands of individual public school teachers, by name, derived from a statistical model of how much their students’ test scores had risen. Measurement specialists objected immediately, and the objection was not that the model was worthless. It was that the ratings had been printed as points. Each teacher’s estimate carried an uncertainty wide enough that many of the people sorted into different categories were statistically indistinguishable from one another, and none of that width made it into the ranking a parent read over breakfast.

That is the shape of the failure, and it is a failure of construction rather than of humility. An interval is not commentary on a score but part of it, so a point estimate released without one is not a cautious claim made quietly, it is an incomplete claim made confidently, and everyone downstream treats the missing width as zero. The number gets copied into a spreadsheet, then into a decision, then into a sentence about a person, and at each hop the doubt that never travelled with it is assumed not to exist.

The reassurance most instruments offer instead is a single reliability coefficient, usually Cronbach’s alpha, treated as a certificate. Klaas Sijtsma’s 2009 paper in Psychometrika is the standard citation for why it is not one. Alpha is a lower bound rather than an estimate of reliability, it rises simply because you added more items, and a high value says nothing about whether those items measure one construct or five. It is possible to build a questionnaire with an impressive alpha, no coherent thing being measured, and a page of scores that look precise to three significant figures.

The honest version of the same classical machinery does the opposite. It converts reliability into a standard error of measurement, and from that into a band around every score. Not your resilience is 62, but your resilience is near 62, give or take about 8, and that range is what the instrument can stand behind. Read plainly, the band says something a bare number cannot: here is the size of my own ignorance about you. The professional standards for educational and psychological testing have required exactly this reporting for years, which is worth remembering the next time a product ships an integer as though the requirement were paperwork.

The discipline gets sharper once you ask what the band is for. It is not decoration and it is not liability management. It is the thing that determines whether a comparison is allowed at all. Two scores whose intervals overlap have not been shown to differ, which means a ranking built from point estimates can order people the underlying evidence cannot order. It is also the thing that determines whether change over time is real: a person’s score moving from 62 to 68 means nothing if the instrument’s error is 8, and the disciplined system says so rather than congratulating anyone. Our own research notes reached the sentence we have not improved on. Scores without confidence decay into superstition, and the decay is fast, because a number that has shed its uncertainty is indistinguishable from a fact.

None of this argues for silence. It argues for a small change in what counts as a complete output, one that any team can make on a Tuesday: a score arriving without its width is a bug report, not a result. Wire the interval into the record beside the value, and the whole downstream chain inherits an honesty nobody has to remember to add. The width was never the caveat on the finding. It was half of what you found.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 4
  1. Klaas Sijtsma (2009). On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha (Psychometrika 74(1), 107-120)

    The load-bearing citation for the brick's second move: alpha is a lower bound, it climbs with item count, and it does not establish that a scale measures one construct.

    Comment on this source
  2. American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing

    The requirement that reported scores be accompanied by information about measurement error, which is why a bare integer is a departure from professional practice rather than a stylistic choice.

    Comment on this source
  3. Los Angeles Times (2010). Publication of value-added effectiveness ratings for individual Los Angeles Unified teachers

    The brick's named, dated case. The dispute was not primarily about whether the model measured something but about publishing point estimates whose intervals were wide enough to blur the categories readers drew from them.

    Comment on this source
  4. Derek Briggs and Ben Domingue (2011). Due Diligence and the Evaluation of Teachers (National Education Policy Center review of the Los Angeles value-added analysis)

    The re-analysis behind the instability grading: alternative model specifications moved a substantial share of teachers across category boundaries, which is the practical face of an unreported interval.

    Comment on this source
Claims and confidence 5
  1. verified

    Cronbach's alpha is a lower bound that rises with the number of items and does not establish that a scale measures one construct.

    Sijtsma 2009 and the subsequent reliability literature; uncontroversial in psychometrics.

    Respond to this claim
  2. verified

    Professional testing standards require reported scores to be accompanied by information about measurement error.

    Standards for Educational and Psychological Testing (2014), reliability and score-reporting chapters.

    Respond to this claim
  3. verified

    In August 2010 the Los Angeles Times published value-added effectiveness ratings for thousands of individual public school teachers by name.

    The publication itself and contemporaneous coverage of it.

    Respond to this claim
  4. directional

    The published value-added ratings carried uncertainty wide enough that teachers sorted into different effectiveness categories were often not statistically distinguishable.

    Directional. Independent re-analyses of the same data reported wide intervals and substantial category movement under alternative specifications; the exact share depends on the model, and no single figure is cited here as settled.

    Respond to this claim
  5. verified

    Two scores whose intervals overlap have not been shown to differ, so a ranking of point estimates can order people the evidence cannot order.

    Property of interval estimation rather than an empirical finding.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to Scores without confidence

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target Scores without confidence

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.