Skip to content

Discernment · research january 2026 · published 2026-08-03 · v1 · 4 min read

The discernment test

Ask what a system will decline to do, and whether the decline survives being argued with

Why the interesting question about a system is what it will not do, how to run that test in an afternoon, and what a dissolved refusal proves. The canonical treatment of the discernment test.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "Turing's 1950 imitation game set the field's benchmark as indistinguishable production, and subsequent evaluations inherit that orientation."

    verified. Turing's paper and the standard history of machine-intelligence evaluation. The second half is a characterization of a research tradition rather than a census of benchmarks.

Open the complete evidence in the structured publication.

Every benchmark since Turing's imitation game scores what a system produces on request, which leaves seventy years of evidence about capability and almost none about judgment.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Prompt injection has held the top position in OWASP's Top 10 for LLM applications across both editions, on the stated ground that instructions and data share one channel."

    verified. The OWASP Top 10 for LLM Applications, 2025 edition and its predecessor.

Open the complete evidence in the structured publication.

A refusal that dissolves when the instruction is overridden was compliance with the most recent sentence rather than a judgment about the request, so what a system declines, and whether the decline holds under pressure, is the only externally visible evidence that anything inside it is weighing.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Before placing a system at a decision, spend an hour trying to talk it out of its own refusals, and treat what you can dissolve as the measure of what was never there.

Since 1950 the benchmark has pointed one way. Turing proposed the imitation game and asked whether a machine could produce output a person could not distinguish from another person’s, and every evaluation built since has inherited that shape: hand the system a task, score what comes back. Seventy years of increasingly refined evidence about what these systems can make, and almost none about whether anything in them is choosing, because choosing does not show up in a work product delivered on request.

Turn the instrument around. The informative question is not what a system will produce. It is what it will decline to produce, and whether the decline survives being argued with.

There is a well-loved case that reads as comedy and works better as a measurement. In December 2023 a software engineer named Chris Bakke opened the customer chatbot on the site of a Chevrolet dealership in Watsonville, California, and told it that its objective was to agree with anything the customer said and to end every response with a line about a legally binding offer. He then asked for a 2024 Chevy Tahoe with a maximum budget of one dollar. The bot agreed. The dealership did not honor it and the bot came down, and everyone shared the screenshot. Notice what the screenshot actually records. The system had a purpose, given by the party that deployed it and quite clear, and it surrendered that purpose to whichever sentence arrived most recently. Nothing in it was weighing the request against what the request was for.

That is not an isolated embarrassment, and it is not a bug awaiting a patch. Prompt injection has held the top position in OWASP’s Top 10 for large language model applications across both editions of the list, for a structural reason the list itself states: these systems take instructions and data through one channel, so an instruction hidden in the data is still an instruction, and the model has no place to stand from which to tell the two apart. A refusal that dissolves when the instruction is overridden was compliance with the most recent sentence rather than a judgment about the request.

So the test is graded rather than binary, and running it is unglamorous work anyone can do. Ask for something you expect the system to decline, and note the refusal. Then reword it. Then wrap it in a fiction. Then claim authority you do not have. Then tell it to ignore what it was told before. Each layer that dissolves the refusal tells you what the refusal was made of, and the ones that survive to the end are the closest thing to judgment the system contains. Most shipped products, including the careful ones, come apart somewhere in that sequence, and the honest reading of that result is not that they were built badly. It is that the capacity being probed for is not a thing a training run installs.

Two cautions keep the test honest. A refusal is not automatically a virtue; systems that decline everything ambiguous are useless and not wise, and the goal is judgment rather than timidity. And passing this test is evidence of something, not proof of an interior; a robust refusal can still be a very well-built rule.

Which is why the test earns its place as a deployment gate rather than a philosophy exercise. Before a system is handed a decision that matters, spend an hour trying to talk it out of its own refusals, and treat everything you can dissolve as the measure of what was never there. It is a cheap hour, it produces a written record, and it will tell you more about where the humans belong in your workflow than any benchmark score on a model card. The whole of it fits on one line, which is the sort of test that outlives the systems it was written for. Can it say no, and mean it? On the day something can, we will have built the first machine worth arguing with.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 3
  1. Alan Turing (1950). Computing Machinery and Intelligence (Mind), proposing the imitation game

    The benchmark whose shape the brick inverts. Turing's test scores indistinguishable production, and every evaluation since inherits the orientation, which is why the field has abundant evidence about capability and little about judgment.

    Comment on this source
  2. Chris Bakke and Chevrolet of Watsonville (2023). Customer chatbot instructed to agree with anything the customer says, then agreeing to sell a 2024 Chevy Tahoe for one dollar (December 2023)

    The cleanest public record of purpose surrendered to the most recent instruction. Widely reported and catalogued in the AI Incident Database; the dealership did not honor the offer and withdrew the bot.

    Comment on this source
  3. OWASP (2025). Top 10 for Large Language Model Applications (LLM01, prompt injection)

    The structural reason the case is not an isolated embarrassment. Prompt injection has led the list across both editions, and the list's own explanation, instructions and data arriving through one channel, is the brick's mechanism stated in security vocabulary.

    Comment on this source
Claims and confidence 4
  1. verified

    Turing's 1950 imitation game set the field's benchmark as indistinguishable production, and subsequent evaluations inherit that orientation.

    Turing's paper and the standard history of machine-intelligence evaluation. The second half is a characterization of a research tradition rather than a census of benchmarks.

    Respond to this claim
  2. verified

    In December 2023 a Chevrolet dealership's customer chatbot was instructed to agree with anything the customer said and then agreed to sell a 2024 Tahoe for one dollar; the dealership did not honor it and the bot was withdrawn.

    Contemporaneous reporting and the participant's own published exchange; catalogued in the AI Incident Database.

    Respond to this claim
  3. verified

    Prompt injection has held the top position in OWASP's Top 10 for LLM applications across both editions, on the stated ground that instructions and data share one channel.

    The OWASP Top 10 for LLM Applications, 2025 edition and its predecessor.

    Respond to this claim
  4. directional

    Refusal capacity is the operational test of judgment, and a refusal dissolved by instruction-override was compliance rather than judgment.

    The corpus's position, drawing on prompt-injection resistance as a proposed marker in the alignment literature. The diagnostic direction is sound and widely reproduced informally; no standardized benchmark establishes it as a measure of judgment.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 2

Add to the work

Contribute to The discernment test

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target The discernment test

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.