Skip to content

Article · research july 2026 · published 2026-08-03 · v1 · 3 min read

The crutch effect

Help that bypasses the struggle shows up as improvement now and damage later, and only the later measurement is honest

The measured case for restraint in assistance, from a thousand students whose scores rose while their skills fell. The canonical treatment of the crutch effect.

Topics: Receipts , Deskilling

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Students given unrestricted GPT-4 access improved practice performance by 48 percent and scored 17 percent worse than controls on the subsequent unassisted exam."

    verified. Bastani et al., PNAS 2025, randomized deployment across roughly one thousand high school students; verified against the published paper and the university's own account.

Open the complete evidence in the structured publication.

Assistance that raises performance while present can lower capability once removed, and every metric taken during the assistance reports the opposite of what is happening.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "A tutored variant of the same model, with prompts that withheld answers and scaffolded attempts, largely eliminated the learning harm."

    verified. The GPT Tutor arm of the same study.

Open the complete evidence in the structured publication.

The cognitive struggle the help removes was the thing doing the encoding, so the metric taken with the tool in hand improves exactly as the capability it stands in for erodes, and the only instrument that would catch the harm is the unassisted test nobody runs.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Price any assistive tool by the measurement taken with the tool absent, because the assisted number is a claim about the tool and only the unassisted number is a claim about you.

In 2025, Bastani and colleagues published in PNAS the cleanest measurement yet taken of what unrestricted machine help does to a person’s capability. Nearly a thousand high school mathematics students in Turkey were given access to GPT-4 while practicing. With the model available, their practice performance rose by 48 percent, which is the number a dashboard would celebrate and a parent would pay for. Then the model was taken away and the students sat an ordinary exam. They scored 17 percent worse than classmates who had never had the model at all. The help had not accelerated their learning. It had stood in for their learning, and the substitution was invisible for exactly as long as the help was present.

The mechanism is uncomfortable because it locates the harm inside the benefit. Struggling with a problem is not the unfortunate cost of acquiring a skill; the struggle is the acquisition, the effortful encoding by which a method becomes something you own. Assistance that removes the struggle removes the encoding, while every measurement taken during the assistance improves, since the measurement can no longer tell the difference between what you can do and what you can do accompanied. The two numbers come apart in silence. The assisted metric rises as the capability it stands in for erodes, and the only instrument that would catch the divergence is the unassisted test, which is precisely the test that a person carrying a helpful tool never has a reason to run.

The study’s authors reached for the parallel the aviation industry has documented for decades: pilots who fly highly automated cockpits lose hand-flying proficiency through disuse, which is why the FAA formally urged operators in 2013 to make their pilots fly manually more often. The skill decays across exactly the years before the moment it is abruptly needed, and the autopilot’s reliability is what funds the decay. But the study’s second arm matters as much as its warning. A different group of students used the same model wrapped in tutoring constraints, prompts that withheld answers, forced attempts, and scaffolded the struggle instead of replacing it, and the harm largely disappeared. Same capability, different shape, opposite outcome. The damage was never a property of what the model could do. It was a property of what the product let it do, which makes this the measured case for below-threshold design: capability deliberately withheld is not capability wasted.

So the move is a discipline of measurement. Whatever the tool, price it by the test taken with the tool absent, because the assisted number is a claim about the tool and only the unassisted number is a claim about you. The products that deserve your trust are the ones willing to be graded that way, and the capability worth building is the kind that is still there when the help is gone.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 2
  1. Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman (2025). Generative AI without guardrails can harm learning, evidence from high school mathematics (PNAS)

    The measured case: roughly a thousand Turkish high school students, a 48 percent rise in assisted practice performance and a 17 percent fall on the subsequent unassisted exam, with a guardrailed tutor arm that largely eliminated the harm.

    Comment on this source
  2. Federal Aviation Administration (2013). Safety Alert for Operators 13002, Manual Flight Operations

    The aviation precedent the study's authors invoke: autoflight dependency degrades hand-flying proficiency, and the regulator formally urged more manual flight because of it.

    Comment on this source
Claims and confidence 3
  1. verified

    Students given unrestricted GPT-4 access improved practice performance by 48 percent and scored 17 percent worse than controls on the subsequent unassisted exam.

    Bastani et al., PNAS 2025, randomized deployment across roughly one thousand high school students; verified against the published paper and the university's own account.

    Respond to this claim
  2. verified

    A tutored variant of the same model, with prompts that withheld answers and scaffolded attempts, largely eliminated the learning harm.

    The GPT Tutor arm of the same study.

    Respond to this claim
  3. verified

    Overreliance on cockpit automation degrades manual flying skill, and the decay concentrates in the years before the moment the skill is abruptly required.

    FAA SAFO 13002 and the human-factors literature on automation dependency; the study's authors draw the same parallel.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 2

Add to the work

Contribute to The crutch effect

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The crutch effect

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.