Article · research july 2026 · published 2026-08-03 · v1 · 3 min read
The crutch effect
Help that bypasses the struggle shows up as improvement now and damage later, and only the later measurement is honest
The measured case for restraint in assistance, from a thousand students whose scores rose while their skills fell. The canonical treatment of the crutch effect.
In 2025, Bastani and colleagues published in PNAS the cleanest measurement yet taken of what unrestricted machine help does to a person’s capability. Nearly a thousand high school mathematics students in Turkey were given access to GPT-4 while practicing. With the model available, their practice performance rose by 48 percent, which is the number a dashboard would celebrate and a parent would pay for. Then the model was taken away and the students sat an ordinary exam. They scored 17 percent worse than classmates who had never had the model at all. The help had not accelerated their learning. It had stood in for their learning, and the substitution was invisible for exactly as long as the help was present.
The mechanism is uncomfortable because it locates the harm inside the benefit. Struggling with a problem is not the unfortunate cost of acquiring a skill; the struggle is the acquisition, the effortful encoding by which a method becomes something you own. Assistance that removes the struggle removes the encoding, while every measurement taken during the assistance improves, since the measurement can no longer tell the difference between what you can do and what you can do accompanied. The two numbers come apart in silence. The assisted metric rises as the capability it stands in for erodes, and the only instrument that would catch the divergence is the unassisted test, which is precisely the test that a person carrying a helpful tool never has a reason to run.
The study’s authors reached for the parallel the aviation industry has documented for decades: pilots who fly highly automated cockpits lose hand-flying proficiency through disuse, which is why the FAA formally urged operators in 2013 to make their pilots fly manually more often. The skill decays across exactly the years before the moment it is abruptly needed, and the autopilot’s reliability is what funds the decay. But the study’s second arm matters as much as its warning. A different group of students used the same model wrapped in tutoring constraints, prompts that withheld answers, forced attempts, and scaffolded the struggle instead of replacing it, and the harm largely disappeared. Same capability, different shape, opposite outcome. The damage was never a property of what the model could do. It was a property of what the product let it do, which makes this the measured case for below-threshold design: capability deliberately withheld is not capability wasted.
So the move is a discipline of measurement. Whatever the tool, price it by the test taken with the tool absent, because the assisted number is a claim about the tool and only the unassisted number is a claim about you. The products that deserve your trust are the ones willing to be graded that way, and the capability worth building is the kind that is still there when the help is gone.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 2
-
Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman (2025). Generative AI without guardrails can harm learning, evidence from high school mathematics (PNAS)
The measured case: roughly a thousand Turkish high school students, a 48 percent rise in assisted practice performance and a 17 percent fall on the subsequent unassisted exam, with a guardrailed tutor arm that largely eliminated the harm.
Comment on this source -
Federal Aviation Administration (2013). Safety Alert for Operators 13002, Manual Flight Operations
The aviation precedent the study's authors invoke: autoflight dependency degrades hand-flying proficiency, and the regulator formally urged more manual flight because of it.
Comment on this source
Claims and confidence 3
- verified
Students given unrestricted GPT-4 access improved practice performance by 48 percent and scored 17 percent worse than controls on the subsequent unassisted exam.
Bastani et al., PNAS 2025, randomized deployment across roughly one thousand high school students; verified against the published paper and the university's own account.
Respond to this claim - verified
A tutored variant of the same model, with prompts that withheld answers and scaffolded attempts, largely eliminated the learning harm.
The GPT Tutor arm of the same study.
Respond to this claim - verified
Overreliance on cockpit automation degrades manual flying skill, and the decay concentrates in the years before the moment the skill is abruptly required.
FAA SAFO 13002 and the human-factors literature on automation dependency; the study's authors draw the same parallel.
Respond to this claim
Read next
-
Receipts · read
Restraint has now been argued from both ends of the room, so follow it to the money, where the incentive to keep you coming back turns out to be a published term in somebody's pay.
The return rate bonus
-
Deskilling · read
The student's version of the harm has an expert's version, and it does not require the machine to be wrong; being right is the mechanism.
The competence ceiling