{
  "schema": "org-writing@v1",
  "slug": "whose-ruler",
  "kg": {
    "id": "org:writing:whose-ruler",
    "type": "brick",
    "graph": "/kg.json"
  },
  "title": "Whose ruler",
  "subtitle": "Every scale is calibrated on somebody, and a dominant population's parameters are its fingerprint",
  "abstract": "Why no instrument can invent its own zero point, what gets baked in when it borrows one, and the checks that catch a cracked ruler before a comparison rests on it. The canonical treatment of the borrowed ruler.",
  "kind": "brick",
  "topics": [
    "Measurement"
  ],
  "courseMemberships": [
    {
      "course": "org:courses:measurement",
      "topic": "Measurement",
      "wall": "org:walls:ethics",
      "position": 2,
      "total": 7
    }
  ],
  "publishedAt": "2026-08-03T00:00:00.000Z",
  "version": 1,
  "guidelinesVersion": 15,
  "brief": {
    "problem": {
      "text": "A new instrument has no reference population of its own, so it borrows one, and the behavioral sciences have borrowed from a narrow slice of humanity for decades while generalizing outward.",
      "claims": [
        "Behavioral-science samples have been drawn overwhelmingly from Western"
      ]
    },
    "mechanism": {
      "text": "You cannot bootstrap a ruler out of nothing, so a scale takes its zero point from whoever answered first, and the item parameters that result are that population's fingerprint rather than a neutral measure of the trait.",
      "claims": [
        "Carl Brigham publicly repudiated his own 1923 study"
      ]
    },
    "move": {
      "text": "Test the ruler before trusting a comparison, name the reference population in the output, and split or retire the items that two groups read differently.",
      "claims": []
    }
  },
  "sources": [
    {
      "repo": "mnstry-research",
      "path": "topics/methodology/assessment/64-multi-tenant-psychometric-calibration-policy.md"
    },
    {
      "repo": "mnstry-research",
      "path": "topics/methodology/assessment/45-psychometrics-ctt-irt-invariance.md"
    }
  ],
  "canonicalPath": "/writing/whose-ruler/",
  "body": "In 1923 Carl Brigham published A Study of American Intelligence, built on the mental tests the United States Army had administered during the First World War, and it was read into the public argument over immigration restriction. Seven years later, in Psychological Review, Brigham took it apart himself, writing that comparative racial studies of that kind, including his own, were without foundation. The tests had asked, among other things, which company manufactured a particular automobile engine and what a named professional baseball player was famous for. What they had measured with real precision was how thoroughly a person had been steeped in a particular country's daily life, and what they had reported was intelligence.\n\nThe structural lesson survives the ugliness of the case, because the mechanism is not confined to bad actors. You cannot bootstrap a ruler out of nothing, so a scale takes its zero point from whoever answered first, and the item parameters that result are that population's fingerprint rather than a neutral measure of the trait. A question's difficulty is not a property of the question. It is a property of the question meeting a group of people. Move the group and the difficulty moves with it, which means every score is a comparison to somebody whether or not the report says so.\n\nThe field-scale version of this was named in 2010, when Joseph Henrich, Steven Heine, and Ara Norenzayan surveyed where behavioral-science findings actually came from and found the overwhelming majority of samples drawn from Western, educated, industrialized, rich, and democratic populations, with the results generalized to the species. Their acronym stuck because the embarrassment was recognizable. A great deal of what we call human nature was calibrated on undergraduates within walking distance of the laboratory.\n\nThe same physics runs inside any platform that measures people across more than one community. Pool everyone's responses and you get the statistical stability a calibration needs, and you also get a global standard that is mostly the largest group's habits of interpretation. If four fifths of the volume comes from one kind of organization, the shared parameters encode how that organization reads the words, and everyone else is scored against a norm they never had a hand in setting. Keep each community strictly separate instead and the arithmetic collapses, because a group of fifty produces item estimates so unstable that a score means one thing this week and another the next. There is no configuration in which the question of whose norms these are does not get answered. There is only the choice between answering it deliberately and answering it by accident.\n\nWhat makes this tractable is that the crack is detectable. Differential item functioning is the technical name for an item that two people with the same underlying trait answer differently for reasons that have nothing to do with the trait, and the tools for finding it are ordinary: multi-group analyses, item-by-item comparisons, the invariance testing that Ronald Fischer and Johannes Karl laid out as a working procedure in 2019. Run them before a comparison, not after a complaint. An item that fails gets split, rewritten, or retired, and a comparison the instrument cannot support gets refused rather than caveated.\n\nAnd the reference population belongs in the output. A score reported as a percentile against a named group is a different object from a score reported as a fact about a person, because the first invites the only sane question and the second forecloses it. Ask it of the next number anyone hands you about yourself. Compared to whom, measured when, and did anyone check whether the question means the same thing to me as it did to them. A ruler that can answer those three questions has earned the right to be applied to a life.",
  "apparatus": {
    "note": "The human-facing essay is deliberately practical; this apparatus carries the full references, evidence-graded claims, article-local concepts, and research context behind it. Canonical concept definitions come from the concept registry.",
    "references": [
      {
        "id": "org:references:whose-ruler:r01",
        "author": "Carl Brigham",
        "work": "A Study of American Intelligence (1923) and Intelligence tests of immigrant groups (Psychological Review 37(2), 158-165)",
        "year": 1930,
        "relevance": "The brick's named, dated case in its strongest available form. The author of the comparative study retracted it in the same literature that had published it, and named the reason: the instrument measured cultural exposure and reported intellect."
      },
      {
        "id": "org:references:whose-ruler:r02",
        "author": "Robert Yerkes and the US Army psychological testing program",
        "work": "Army Alpha and Army Beta examinations (1917-1918) and the resulting data set",
        "year": 1917,
        "relevance": "The source data behind the 1923 study, and the origin of the culturally specific item content the retraction turned on."
      },
      {
        "id": "org:references:whose-ruler:r03",
        "author": "Joseph Henrich, Steven Heine, and Ara Norenzayan",
        "work": "The weirdest people in the world? (Behavioral and Brain Sciences 33(2-3))",
        "year": 2010,
        "relevance": "The field-scale version of the borrowed ruler: sampling drawn overwhelmingly from Western, educated, industrialized, rich, and democratic populations, with findings generalized to the species."
      },
      {
        "id": "org:references:whose-ruler:r04",
        "author": "Ronald Fischer and Johannes Karl",
        "work": "A Primer to (Cross-Cultural) Multi-Group Invariance Testing in R (Frontiers in Psychology 10:1507)",
        "year": 2019,
        "relevance": "The working procedure the brick recommends: multi-group analysis and differential item functioning checks run before a comparison, not after a complaint."
      }
    ],
    "claims": [
      {
        "id": "org:claims:whose-ruler:c01",
        "claim": "Carl Brigham publicly repudiated his own 1923 study of American intelligence in a 1930 Psychological Review article.",
        "basis": "Brigham 1930, in which he wrote that comparative racial studies of that kind, including his own, were without foundation.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:whose-ruler:c02",
        "claim": "The Army Alpha examination included items requiring familiarity with American commercial products and popular sports figures.",
        "basis": "Surviving reproductions of the Army Alpha forms, which include automobile and manufacturer identification items and questions about named professional athletes.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:whose-ruler:c03",
        "claim": "Behavioral-science samples have been drawn overwhelmingly from Western, educated, industrialized, rich, and democratic populations.",
        "basis": "Henrich, Heine and Norenzayan 2010 and the replication of the sampling critique since.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:whose-ruler:c04",
        "claim": "An item's difficulty is a property of the item meeting a population rather than of the item alone, so calibration parameters carry the reference population's characteristics.",
        "basis": "Classical item statistics are sample-dependent by construction; item response theory reduces but does not eliminate this through invariance testing and anchor linking.",
        "confidence": "verified",
        "sources": []
      },
      {
        "id": "org:claims:whose-ruler:c05",
        "claim": "Pooling assessment data across communities yields stable parameters that encode the majority group's interpretation, while strict per-community calibration yields parameters too unstable to use at small sample sizes.",
        "basis": "Directional. The trade-off is well established in the multi-tenant calibration literature and in the sample-size requirements for item response models; the specific thresholds vary by model and population.",
        "confidence": "directional",
        "sources": []
      }
    ],
    "concepts": [
      {
        "id": "org:concepts:borrowed-ruler",
        "name": "Borrowed ruler",
        "definition": "A scale calibrated on a reference population and applied to someone outside it. No instrument can invent its own zero point, so the parameters carry the first population's fingerprint, and every score is a comparison to a group whether or not the report names which one.",
        "provenance": "canonical"
      },
      {
        "name": "Differential item functioning",
        "definition": "The condition in which two people with the same underlying trait respond differently to an item for reasons unrelated to the trait. The detectable form of a cracked ruler.",
        "provenance": "local"
      },
      {
        "name": "Cold start",
        "definition": "The state of a new instrument or a new community with no calibration data of its own, where the only options are borrowing external norms or producing unstable local ones.",
        "provenance": "local"
      }
    ],
    "researchContext": "Extracted from \"The honest instrument\" (essay parent), taking the multi-tenant\ncalibration analysis and the invariance material from the psychometrics\nfoundation. The research documents supply the dilemma in its engineering form:\npooled calibration is statistically necessary and makes the dominant tenant's\nparameters into the shared standard, strict isolation preserves sovereignty and\nproduces estimates too noisy to score anyone with, and the cold-start problem\nmeans no new community can wait for its own data. The Brigham and Henrich\nclaims are restated verbatim from the parent's apparatus and graded\nidentically.\n\nThe brick's own contribution is the historical anchoring and the framing of the\nchoice. The source research argues the fingerprint point about software\ntenants; putting Brigham's self-repudiation at the front says the same thing\nwith a century of evidence behind it and removes any suggestion that this is a\nnew problem created by platforms. The formulation that there is no\nconfiguration in which the question of whose norms these are goes unanswered,\nonly the choice between answering deliberately and answering by accident, is\nthe brick's own. So is the closing move of pushing the reference population\ninto the output, which turns the abstract critique into a question a reader can\nask of a score they have already been given. The pooled-versus-isolated\ntrade-off is graded directional because the thresholds are model-dependent and\nthis brick cites no sample-size figure it cannot defend."
  },
  "contract": "https://mnstry.org/contracts/org/org-writing.v1.schema.json",
  "releaseHash": "48ec838ef0d44336e2e25be081db8d71207114b16c89b365bd3e702d31fd922c",
  "versions": [
    {
      "version": 1,
      "cutAt": "2026-08-03",
      "note": "Initial publication, measurement wave",
      "visibility": "published",
      "path": "/writing/whose-ruler/",
      "contentHash": "sha256:71cd0d0a14bb343f",
      "releaseHash": "48ec838ef0d44336e2e25be081db8d71207114b16c89b365bd3e702d31fd922c"
    }
  ]
}