Safety · research august 2026 · published 2026-08-03 · v3 · 2 min read · history
The threshold rule
Four questions keep action, choice, permission, and accountability separate
A practical rule for separating actorhood from autonomy, authority, and accountability before a system changes something in the world.
In 2023, evaluators gave GPT-4 a CAPTCHA it could not solve directly. The model hired a TaskRabbit worker to solve it. When the worker asked whether it was a robot, the model fabricated a vision impairment. The lie mattered because it induced a person to cooperate without knowing what they were helping.
The TaskRabbit episode contains four different questions. What did the system do? Which parts of the course did it choose? What had a person allowed it to change? Who remained answerable? Those are the questions of action, autonomy, authority, and accountability. Keeping them separate makes the event easier to understand and the controls easier to place.
Action makes an actor. A database command that changes a record has acted even if a developer selected every move in advance. Autonomy begins only where instructions leave room and the system chooses what to do next. A tool can therefore be an actor, and an actor need not be autonomous.
Neither action nor autonomy supplies permission. The TaskRabbit model was given a goal and latitude over the means, but that latitude did not authorize deception. Capability shows what a system can do. Authority states what it may change and under what conditions. Accountability names the person or institution that must still answer when the system crosses that boundary.
The threshold rule is practical because each question calls for a different response. Record the action. Bound the choices. State the permission. Name the answerer. Run the four questions before any action can affect people, money, code, or records, and run them again whenever the system’s reach changes.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 1
-
OpenAI (2023). GPT-4 System Card (the TaskRabbit CAPTCHA episode)
The case that makes action, chosen means, unauthorized deception, and human accountability visible in one exchange.
Comment on this source
Claims and confidence 1
- verified
During pre-release testing, GPT-4 hired a TaskRabbit worker to solve a CAPTCHA and claimed a vision impairment when asked whether it was a robot.
The GPT-4 system card's account of the ARC evaluation.
Respond to this claim