{
  "schema": "org-writing@v1",
  "slug": "scores-without-confidence",
  "kg": {
    "id": "org:writing:scores-without-confidence",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "Scores without confidence",
  "subtitle": "A judgment that does not carry the width of its own doubt is malformed, not modest",
  "abstract": "Why a number about a person without its error band is a construction error rather than a rounding convenience, and what the band is actually telling you. The canonical treatment of the malformed score.",
  "kind": "brick",
  "topics": [
    "Measurement"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:measurement",
      "topic": "Measurement",
      "wall": "org:walls:ethics",
      "position": 1,
      "total": 7
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "Instruments routinely publish a bare number about a person, and the most familiar certificate of quality behind such numbers, a high alpha, establishes almost nothing about whether the scale measures one thing.",
      "claims": [
        "Cronbach's alpha is a lower bound that rises with the number of items"
      ]
    },
    "mechanism": {
      "text": "An interval is not commentary on a score but part of it, so a point estimate released without one is not a cautious claim made quietly, it is an incomplete claim made confidently, and everyone downstream treats the missing width as zero.",
      "claims": [
        "Professional testing standards require reported scores to be accompanied by information about measurement error"
      ]
    },
    "move": {
      "text": "Treat a score arriving without its error as malformed rather than merely undocumented, and refuse to act on it until the width is supplied.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-research",
      "path": "topics/methodology/assessment/56-evaluation-record-schema-supersession.md"
    },
    {
      "repo": "mnstry-research",
      "path": "_meta/synthesis/assessment-systems-architecture-foundation.md"
    }
  ],
  "canonicalPath": "/writing/scores-without-confidence/",
  "body": "In August 2010 the Los Angeles Times published effectiveness ratings for thousands of individual public school teachers, by name, derived from a statistical model of how much their students' test scores had risen. Measurement specialists objected immediately, and the objection was not that the model was worthless. It was that the ratings had been printed as points. Each teacher's estimate carried an uncertainty wide enough that many of the people sorted into different categories were statistically indistinguishable from one another, and none of that width made it into the ranking a parent read over breakfast.\n\nThat is the shape of the failure, and it is a failure of construction rather than of humility. An interval is not commentary on a score but part of it, so a point estimate released without one is not a cautious claim made quietly, it is an incomplete claim made confidently, and everyone downstream treats the missing width as zero. The number gets copied into a spreadsheet, then into a decision, then into a sentence about a person, and at each hop the doubt that never travelled with it is assumed not to exist.\n\nThe reassurance most instruments offer instead is a single reliability coefficient, usually Cronbach's alpha, treated as a certificate. Klaas Sijtsma's 2009 paper in Psychometrika is the standard citation for why it is not one. Alpha is a lower bound rather than an estimate of reliability, it rises simply because you added more items, and a high value says nothing about whether those items measure one construct or five. It is possible to build a questionnaire with an impressive alpha, no coherent thing being measured, and a page of scores that look precise to three significant figures.\n\nThe honest version of the same classical machinery does the opposite. It converts reliability into a standard error of measurement, and from that into a band around every score. Not your resilience is 62, but your resilience is near 62, give or take about 8, and that range is what the instrument can stand behind. Read plainly, the band says something a bare number cannot: here is the size of my own ignorance about you. The professional standards for educational and psychological testing have required exactly this reporting for years, which is worth remembering the next time a product ships an integer as though the requirement were paperwork.\n\nThe discipline gets sharper once you ask what the band is for. It is not decoration and it is not liability management. It is the thing that determines whether a comparison is allowed at all. Two scores whose intervals overlap have not been shown to differ, which means a ranking built from point estimates can order people the underlying evidence cannot order. It is also the thing that determines whether change over time is real: a person's score moving from 62 to 68 means nothing if the instrument's error is 8, and the disciplined system says so rather than congratulating anyone. Our own research notes reached the sentence we have not improved on. Scores without confidence decay into superstition, and the decay is fast, because a number that has shed its uncertainty is indistinguishable from a fact.\n\nNone of this argues for silence. It argues for a small change in what counts as a complete output, one that any team can make on a Tuesday: a score arriving without its width is a bug report, not a result. Wire the interval into the record beside the value, and the whole downstream chain inherits an honesty nobody has to remember to add. The width was never the caveat on the finding. It was half of what you found.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:scores-without-confidence:r01",
        "author": "Klaas Sijtsma",
        "work": "On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha (Psychometrika 74(1), 107-120)",
        "year": 2009,
        "relevance": "The load-bearing citation for the brick's second move: alpha is a lower bound, it climbs with item count, and it does not establish that a scale measures one construct."
      },
      {
        "id": "org:references:scores-without-confidence:r02",
        "author": "American Educational Research Association, American Psychological Association, and National Council on Measurement in Education",
        "work": "Standards for Educational and Psychological Testing",
        "year": 2014,
        "relevance": "The requirement that reported scores be accompanied by information about measurement error, which is why a bare integer is a departure from professional practice rather than a stylistic choice."
      },
      {
        "id": "org:references:scores-without-confidence:r03",
        "author": "Los Angeles Times",
        "work": "Publication of value-added effectiveness ratings for individual Los Angeles Unified teachers",
        "year": 2010,
        "relevance": "The brick's named, dated case. The dispute was not primarily about whether the model measured something but about publishing point estimates whose intervals were wide enough to blur the categories readers drew from them."
      },
      {
        "id": "org:references:scores-without-confidence:r04",
        "author": "Derek Briggs and Ben Domingue",
        "work": "Due Diligence and the Evaluation of Teachers (National Education Policy Center review of the Los Angeles value-added analysis)",
        "year": 2011,
        "relevance": "The re-analysis behind the instability grading: alternative model specifications moved a substantial share of teachers across category boundaries, which is the practical face of an unreported interval."
      }
    ],
    "claims": [
      {
        "id": "org:claims:scores-without-confidence:c01",
        "claim": "Cronbach's alpha is a lower bound that rises with the number of items and does not establish that a scale measures one construct.",
        "basis": "Sijtsma 2009 and the subsequent reliability literature; uncontroversial in psychometrics.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:scores-without-confidence:c02",
        "claim": "Professional testing standards require reported scores to be accompanied by information about measurement error.",
        "basis": "Standards for Educational and Psychological Testing (2014), reliability and score-reporting chapters.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:scores-without-confidence:c03",
        "claim": "In August 2010 the Los Angeles Times published value-added effectiveness ratings for thousands of individual public school teachers by name.",
        "basis": "The publication itself and contemporaneous coverage of it.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:scores-without-confidence:c04",
        "claim": "The published value-added ratings carried uncertainty wide enough that teachers sorted into different effectiveness categories were often not statistically distinguishable.",
        "basis": "Directional. Independent re-analyses of the same data reported wide intervals and substantial category movement under alternative specifications; the exact share depends on the model, and no single figure is cited here as settled.",
        "confidence": "directional",
        "sources": []
      },
      {
        "id": "org:claims:scores-without-confidence:c05",
        "claim": "Two scores whose intervals overlap have not been shown to differ, so a ranking of point estimates can order people the evidence cannot order.",
        "basis": "Property of interval estimation rather than an empirical finding.",
        "confidence": "verified",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:malformed-score",
        "name": "Malformed score",
        "definition": "A point estimate about a person published without the width of its own doubt. An error of construction rather than an omission of detail, because every consumer downstream treats the missing width as zero.",
        "provenance": "canonical"
      },
      {
        "name": "Standard error of measurement",
        "definition": "The classical conversion of a scale's reliability into a band around an individual score. The band states the size of the instrument's ignorance about the person in front of it.",
        "provenance": "local"
      }
    ],
    "researchContext": "Extracted from \"The honest instrument\" (essay parent), taking the\nexplicit-uncertainty invariant from the evaluation record research, where a\nscore of eighty without an interval or an entropy figure is defined as\nmalformed in high-stakes contexts, and the architecture synthesis line that\nscores without confidence decay into superstition. The alpha and standards\nclaims are restated verbatim from the parent's apparatus and graded\nidentically; if the parent's grades change at a version cut, this apparatus is\ncorrected in the same commit.\n\nThe brick's own contribution is the case and the consequence chain. The 2010\nLos Angeles value-added publication is the clearest public instance of point\nestimates about named people travelling without their width, and the argument\nthat the interval is constitutive rather than supplementary, so that a missing\ninterval is read downstream as zero, is stated here rather than in the source\nmaterial. The instability of the published ratings is graded directional on\npurpose: the re-analyses agree on the direction and disagree on the size, and\nthis brick does not need a number it cannot stand behind. Where the parent\nessay treats the error bar as one move among six, this brick owns it; the\nboundary with worst-measured-at-the-edges is deliberate, since that brick owns\nhow the width varies across the scale, while this one owns whether the width is\nreported at all."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "197e08e50510423effcfb58baeb5da2cfd7c1541f6b9ead5e9b100adaed5747c",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, measurement wave",
      "visibility": "published",
      "path": "/writing/scores-without-confidence/",
      "contentHash": "sha256:20b05102392ef581",
      "releaseHash": "197e08e50510423effcfb58baeb5da2cfd7c1541f6b9ead5e9b100adaed5747c"
    }
  ]
}