Skip to content

Measurement · research february to may 2026 · published 2026-08-03 · v1 · 3 min read

Whose ruler

Every scale is calibrated on somebody, and a dominant population's parameters are its fingerprint

Why no instrument can invent its own zero point, what gets baked in when it borrows one, and the checks that catch a cracked ruler before a comparison rests on it. The canonical treatment of the borrowed ruler.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Behavioral-science samples have been drawn overwhelmingly from Western, educated, industrialized, rich, and democratic populations."

    verified. Henrich, Heine and Norenzayan 2010 and the replication of the sampling critique since.

Open the complete evidence in the structured publication.

A new instrument has no reference population of its own, so it borrows one, and the behavioral sciences have borrowed from a narrow slice of humanity for decades while generalizing outward.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Carl Brigham publicly repudiated his own 1923 study of American intelligence in a 1930 Psychological Review article."

    verified. Brigham 1930, in which he wrote that comparative racial studies of that kind, including his own, were without foundation.

Open the complete evidence in the structured publication.

You cannot bootstrap a ruler out of nothing, so a scale takes its zero point from whoever answered first, and the item parameters that result are that population's fingerprint rather than a neutral measure of the trait.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Test the ruler before trusting a comparison, name the reference population in the output, and split or retire the items that two groups read differently.

In 1923 Carl Brigham published A Study of American Intelligence, built on the mental tests the United States Army had administered during the First World War, and it was read into the public argument over immigration restriction. Seven years later, in Psychological Review, Brigham took it apart himself, writing that comparative racial studies of that kind, including his own, were without foundation. The tests had asked, among other things, which company manufactured a particular automobile engine and what a named professional baseball player was famous for. What they had measured with real precision was how thoroughly a person had been steeped in a particular country’s daily life, and what they had reported was intelligence.

The structural lesson survives the ugliness of the case, because the mechanism is not confined to bad actors. You cannot bootstrap a ruler out of nothing, so a scale takes its zero point from whoever answered first, and the item parameters that result are that population’s fingerprint rather than a neutral measure of the trait. A question’s difficulty is not a property of the question. It is a property of the question meeting a group of people. Move the group and the difficulty moves with it, which means every score is a comparison to somebody whether or not the report says so.

The field-scale version of this was named in 2010, when Joseph Henrich, Steven Heine, and Ara Norenzayan surveyed where behavioral-science findings actually came from and found the overwhelming majority of samples drawn from Western, educated, industrialized, rich, and democratic populations, with the results generalized to the species. Their acronym stuck because the embarrassment was recognizable. A great deal of what we call human nature was calibrated on undergraduates within walking distance of the laboratory.

The same physics runs inside any platform that measures people across more than one community. Pool everyone’s responses and you get the statistical stability a calibration needs, and you also get a global standard that is mostly the largest group’s habits of interpretation. If four fifths of the volume comes from one kind of organization, the shared parameters encode how that organization reads the words, and everyone else is scored against a norm they never had a hand in setting. Keep each community strictly separate instead and the arithmetic collapses, because a group of fifty produces item estimates so unstable that a score means one thing this week and another the next. There is no configuration in which the question of whose norms these are does not get answered. There is only the choice between answering it deliberately and answering it by accident.

What makes this tractable is that the crack is detectable. Differential item functioning is the technical name for an item that two people with the same underlying trait answer differently for reasons that have nothing to do with the trait, and the tools for finding it are ordinary: multi-group analyses, item-by-item comparisons, the invariance testing that Ronald Fischer and Johannes Karl laid out as a working procedure in 2019. Run them before a comparison, not after a complaint. An item that fails gets split, rewritten, or retired, and a comparison the instrument cannot support gets refused rather than caveated.

And the reference population belongs in the output. A score reported as a percentile against a named group is a different object from a score reported as a fact about a person, because the first invites the only sane question and the second forecloses it. Ask it of the next number anyone hands you about yourself. Compared to whom, measured when, and did anyone check whether the question means the same thing to me as it did to them. A ruler that can answer those three questions has earned the right to be applied to a life.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 4
  1. Carl Brigham (1930). A Study of American Intelligence (1923) and Intelligence tests of immigrant groups (Psychological Review 37(2), 158-165)

    The brick's named, dated case in its strongest available form. The author of the comparative study retracted it in the same literature that had published it, and named the reason: the instrument measured cultural exposure and reported intellect.

    Comment on this source
  2. Robert Yerkes and the US Army psychological testing program (1917). Army Alpha and Army Beta examinations (1917-1918) and the resulting data set

    The source data behind the 1923 study, and the origin of the culturally specific item content the retraction turned on.

    Comment on this source
  3. Joseph Henrich, Steven Heine, and Ara Norenzayan (2010). The weirdest people in the world? (Behavioral and Brain Sciences 33(2-3))

    The field-scale version of the borrowed ruler: sampling drawn overwhelmingly from Western, educated, industrialized, rich, and democratic populations, with findings generalized to the species.

    Comment on this source
  4. Ronald Fischer and Johannes Karl (2019). A Primer to (Cross-Cultural) Multi-Group Invariance Testing in R (Frontiers in Psychology 10:1507)

    The working procedure the brick recommends: multi-group analysis and differential item functioning checks run before a comparison, not after a complaint.

    Comment on this source
Claims and confidence 5
  1. verified

    Carl Brigham publicly repudiated his own 1923 study of American intelligence in a 1930 Psychological Review article.

    Brigham 1930, in which he wrote that comparative racial studies of that kind, including his own, were without foundation.

    Respond to this claim
  2. verified

    The Army Alpha examination included items requiring familiarity with American commercial products and popular sports figures.

    Surviving reproductions of the Army Alpha forms, which include automobile and manufacturer identification items and questions about named professional athletes.

    Respond to this claim
  3. verified

    Behavioral-science samples have been drawn overwhelmingly from Western, educated, industrialized, rich, and democratic populations.

    Henrich, Heine and Norenzayan 2010 and the replication of the sampling critique since.

    Respond to this claim
  4. verified

    An item's difficulty is a property of the item meeting a population rather than of the item alone, so calibration parameters carry the reference population's characteristics.

    Classical item statistics are sample-dependent by construction; item response theory reduces but does not eliminate this through invariance testing and anchor linking.

    Respond to this claim
  5. directional

    Pooling assessment data across communities yields stable parameters that encode the majority group's interpretation, while strict per-community calibration yields parameters too unstable to use at small sample sizes.

    Directional. The trade-off is well established in the multi-tenant calibration literature and in the sample-size requirements for item response models; the specific thresholds vary by model and population.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to Whose ruler

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target Whose ruler

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.