{
  "schema": "org-writing@v1",
  "slug": "worst-measured-at-the-edges",
  "kg": {
    "id": "org:writing:worst-measured-at-the-edges",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "Worst measured at the edges",
  "subtitle": "An instrument is least reliable exactly where its verdicts do the most damage",
  "abstract": "How a questionnaire can be precise about an average person and nearly blind about an exceptional one, and why the average reliability figure conceals it. The canonical treatment of edge blindness.",
  "kind": "brick",
  "topics": [
    "Measurement"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:measurement",
      "topic": "Measurement",
      "wall": "org:walls:ethics",
      "position": 3,
      "total": 7
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "Assessments are built from items pitched at typical people, and the decisions that follow from them are made about atypical ones, so the instrument is asked its hardest questions precisely where it has the least to say.",
      "claims": [
        "Item response theory reports precision that varies by trait level"
      ]
    },
    "mechanism": {
      "text": "Precision depends on where the questions sit, so a scale assembled around the middle of a trait carries its tightest band where nothing is decided and its widest band at the extremes, which is where selection, screening, and alarm all happen.",
      "claims": [
        "Professional testing standards require reported scores to be accompanied by information about measurement error"
      ]
    },
    "move": {
      "text": "Report the error conditionally rather than on average, and refuse a decision the instrument cannot support at that end of the scale even when it supports it comfortably in the middle.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-research",
      "path": "topics/methodology/assessment/45-psychometrics-ctt-irt-invariance.md"
    }
  ],
  "canonicalPath": "/writing/worst-measured-at-the-edges/",
  "body": "Classical test theory hands every person the same error bar, which is a convenient fiction and a consequential one. Item response theory drops the fiction. By modelling each question separately, how hard it is and how sharply it separates someone who has a trait from someone who does not, it can report how much the whole instrument actually knows at every point along the scale. That report is the test information function, and the first time you see one plotted, the shape is a small shock. It is a hill. Information peaks somewhere near the middle of the trait and falls away toward both ends, which means the standard error does the opposite: narrow where the crowd is, wide out where the people are unusual.\n\nPrecision depends on where the questions sit, so a scale assembled around the middle of a trait carries its tightest band where nothing is decided and its widest band at the extremes, which is where selection, screening, and alarm all happen. Nobody convenes a review over an average result. The decisions get made about the person at the top of the distribution and the person at the bottom, and those are exactly the two people the questionnaire was least equipped to describe. An instrument can be entirely sound and still be reporting, in effect, that this candidate is somewhere in the top fifth, give or take a fifth.\n\nConsider what that does to a threshold. A program that admits the top five percent needs the instrument to separate the top five percent from the top fifteen, and at that end of the scale it very often cannot, because there were only ever two or three items difficult enough to discriminate up there. The cutoff still runs. The names still sort. What has actually happened is that a coin flip has been laundered through a decimal point, and the people on both sides of the line have been told something about themselves that the evidence does not contain. The same arithmetic runs at the other end, where a screening tool built to describe ordinary distress is asked to identify the person in danger.\n\nTwo disciplines follow, and neither requires abandoning the instrument. The first is reporting the error conditionally rather than on average. A single reliability coefficient averages the hill flat and publishes the mean as though it applied to everyone; the professional testing standards ask for the error at the score in question, which is the number that actually governs whether this particular verdict is safe. The second is building toward the edges on purpose. Adaptive testing exists in large part for this reason: rather than asking everyone the same forty questions, most of which tell it nothing, the instrument selects the next question to be the most informative one it can ask given what it already believes, and stops when its uncertainty is low enough. Two people finish at different lengths, not because one is better but because the instrument reached honest confidence about one of them sooner. Graduate admissions testing moved to computerized adaptive form in the 1990s on exactly this logic.\n\nWhat the hill really tells you is which people the instrument has not yet earned the right to describe, and that is a gift rather than an embarrassment. A blind spot you can plot is a blind spot you can fix, by writing items hard enough and easy enough to see the ends of the range, or by declining the decision until you have. The alternative is not a better measurement. It is the same ignorance without the map. Any instrument that reports where it stops seeing has told you where to build next, and that is the beginning of an instrument worthy of the people at the edges.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:worst-measured-at-the-edges:r01",
        "author": "Frederic Lord and the item response theory literature",
        "work": "Test information functions and conditional standard errors of measurement",
        "relevance": "The technical core of the brick: information is a function of trait level, peaks where the items are concentrated, and falls away at the extremes, so the standard error is narrowest in the middle and widest at the ends."
      },
      {
        "id": "org:references:worst-measured-at-the-edges:r02",
        "author": "American Educational Research Association, American Psychological Association, and National Council on Measurement in Education",
        "work": "Standards for Educational and Psychological Testing",
        "year": 2014,
        "relevance": "The requirement that measurement error accompany reported scores, and the specific expectation of conditional standard errors where precision varies across the score range. The averaged coefficient is what the standards are written against."
      },
      {
        "id": "org:references:worst-measured-at-the-edges:r03",
        "author": "Rizqy Amelia Zein and Hanif Akhtar",
        "work": "Getting started with the graded response model (International Journal of Psychology)",
        "year": 2024,
        "relevance": "Working treatment of the polytomous model behind per-item discrimination and thresholds, which is where the shape of the information curve comes from on Likert-type instruments."
      },
      {
        "id": "org:references:worst-measured-at-the-edges:r04",
        "author": "Educational Testing Service",
        "work": "Introduction of the computerized adaptive form of the Graduate Record Examinations",
        "year": 1993,
        "relevance": "The named case for adaptivity as a response to edge blindness: selecting the next item to be maximally informative for this examinee rather than administering a fixed form built for the middle."
      }
    ],
    "claims": [
      {
        "id": "org:claims:worst-measured-at-the-edges:c01",
        "claim": "Item response theory reports precision that varies by trait level, so an instrument can be precise in the middle of a scale and nearly uninformative at its extremes.",
        "basis": "Test information functions and conditional standard errors; standard result in the IRT literature.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:worst-measured-at-the-edges:c02",
        "claim": "Professional testing standards require reported scores to be accompanied by information about measurement error.",
        "basis": "Standards for Educational and Psychological Testing (2014), reliability and score-reporting chapters.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:worst-measured-at-the-edges:c03",
        "claim": "A single averaged reliability coefficient conceals the variation of precision across the score range.",
        "basis": "Arithmetic property of averaging a conditional quantity; stated directly in the psychometrics literature on marginal versus conditional reliability.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:worst-measured-at-the-edges:c04",
        "claim": "The Graduate Record Examinations moved to a computerized adaptive form in the 1990s, selecting items by their informativeness for the individual examinee.",
        "basis": "Testing-organization documentation of the transition; the specific adaptive granularity has changed since, and the brick claims only the transition and its rationale.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:worst-measured-at-the-edges:c05",
        "claim": "Adaptive administration reaches a target precision in fewer items than a fixed form of equivalent precision.",
        "basis": "Directional. Consistently reported across the computerized adaptive testing literature, with the size of the reduction dependent on the item bank and the stopping rule.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:edge-blindness",
        "name": "Edge blindness",
        "definition": "The pattern in which an instrument's information peaks near the middle of a trait and falls away at both extremes, so precision is lowest exactly where selection, screening, and alarm decisions are made. Visible in a test information function and hidden by any single averaged reliability figure.",
        "provenance": "canonical"
      },
      {
        "name": "Test information function",
        "definition": "The curve describing how much an instrument knows at each point along a trait, from which the conditional standard error is derived. Its usual shape is a hill centered on the items' difficulty.",
        "provenance": "local"
      },
      {
        "name": "Conditional standard error",
        "definition": "The error at the score in question rather than averaged across the scale. The number that actually governs whether a particular verdict is supportable.",
        "provenance": "local"
      }
    ],
    "researchContext": "Extracted from \"The honest instrument\" (essay parent), taking the item response\ntheory material from the psychometrics foundation research, which states both\nhalves of the brick's argument in engineering register: that item response\ntheory yields per-person standard errors that widen at the extremes, and that a\nscale should not be used to select a top band when its error cannot distinguish\nthat band from a wider one. The precision and standards claims are restated\nverbatim from the parent's apparatus and graded identically.\n\nThe brick's contribution is the inversion that makes the technical fact matter.\nThe research presents varying precision as a reason to prefer item response\ntheory; the brick observes that the variation is not neutral, because the\ndistribution of decisions is the mirror image of the distribution of\ninformation. Nobody convenes a review over an average result. The formulation\nof a threshold decision as a coin flip laundered through a decimal point is the\nbrick's own, as is the closing reading of the information curve as a map of\nwhere to build next rather than a confession. The boundary with\nscores-without-confidence is deliberate and load-bearing: that brick owns\nwhether the width is reported at all, this one owns how the width changes\nacross the scale, and neither re-argues the other's half. The adaptive-testing\nefficiency claim is graded directional because the size of the item saving\ndepends on the bank and the stopping rule, and no figure is cited here."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "d82b0e2056e82642d9077c75f6be05eb664166568f6404df6505f550c8b81eab",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, measurement wave",
      "visibility": "published",
      "path": "/writing/worst-measured-at-the-edges/",
      "contentHash": "sha256:9c2ebb2fa8ad5958",
      "releaseHash": "d82b0e2056e82642d9077c75f6be05eb664166568f6404df6505f550c8b81eab"
    }
  ]
}