Article · research january 2026 · published 2026-08-03 · v2 · 10 min read · history
Discernment
A system can hold every fact and still have no judgment
What the tradition called practical wisdom, why holding more knowledge never produces it, and how to build for the gap without waiting for wise machines. The corpus's document of record on discernment.
Grant the capability case everything it asks for, because it is real. A frontier model has read more than a person reads in forty lifetimes. It holds more of medicine than any physician holds, more of contract law than any attorney, and it will produce a competent draft in either field before you have finished asking for one. Nothing in this essay disputes that. This essay is about the moment immediately after, when all of that holding has to become a choice, and about a gap that more holding does not close.
The gap has a measurable shape. In 2021 a team at the University of Michigan externally validated the sepsis prediction model built into one of the most widely deployed hospital record systems in the United States, across 38,455 hospitalizations. The model identified a third of the patients who went on to develop sepsis and missed the other two thirds, while raising alerts on eighteen percent of everyone admitted. You can read that result as a knowledge failure and conclude the model wants better training data. Or you can read it as what it is, a failure of judgment at the one place in the workflow where judgment was the entire job, and conclude something harder about where a system of that kind belongs in a hospital at all.
The tradition has a word for the missing capacity, and it predates computing by two and a half thousand years. What follows is that word, why the systems we build do not have the thing it names, and what a builder does about it on a Tuesday.
The ladder hides a rung
Data, information, knowledge, wisdom. The staircase is so familiar in knowledge management that it usually gets drawn as a pyramid and then left alone. Russell Ackoff formalized it in 1989 in From Data to Wisdom, and his version carried a fifth term that the popular pyramid quietly dropped: understanding, sitting between knowledge and wisdom. He also offered an estimate of the mind’s composition that reads as a joke until you sit with it, roughly forty percent data, thirty percent information, twenty percent knowledge, ten percent understanding, and virtually no wisdom.
The dropped rung matters because it marks exactly where machine systems stop. A model ascends from data to knowledge by statistical correlation, and it does so magnificently. What it cannot do is the thing the missing rung names, the grasp of why one fact bears on another that lets a person say this is the same situation as that one, wearing different clothes.
But the deeper problem is the picture itself. A ladder implies that each rung is a purer grade of the rung below, that wisdom is knowledge with more of it, and therefore that a system holding enough knowledge arrives at wisdom by accumulation. That is not what the fourth rung is. Wisdom is not knowledge at higher resolution. It is a different faculty operating on a different object, and the ladder metaphor hides the one ingredient it runs on.
Five words for knowing
The Greeks did not have one word for knowing, and the reason to reach for their vocabulary here is not decoration. It is resolution. In the sixth book of the Nicomachean Ethics, Aristotle separates the states of the soul by which we get at truth, and the separation turns out to cut the machine exactly along its seams.
Techne is the rational skill of production, the knowledge of how to bring a thing into being. Machine systems hold it in superhuman measure and hold it value-neutral, which is the whole trouble with it, since the same skill produces the vaccine and the pathogen with equal fluency. Episteme, systematic knowledge of what cannot be otherwise, is simulated well and derivatively; the model knows that force equals mass times acceleration because the sentence occurs in its training corpus, not because it has followed the derivation. Nous, the intuitive grasp of first principles, is absent, which is the old symbol-grounding complaint restated in a better vocabulary. Gnosis, the knowledge that comes only from having undergone something, is absent, and it cannot be otherwise for a thing that has undergone nothing. And phronesis, practical wisdom, the capacity to deliberate well about what is good in a particular situation, is absent for a reason we will come to.
That leaves sophia, theoretical wisdom, which Aristotle builds out of intellect and science together. Here the machine produces something genuinely strange. Ask it for wisdom and it delivers wisdom’s exact shape, cadence, balance, and consolation, assembled out of every wisdom text ever digitized, with nothing behind the shape. We call this zombie sophia, perfect form with an empty interior, and it is the failure mode hardest to catch, because everything you would check is present. What is absent leaves no trace in the output.
The minor premise
Aristotle also supplies the mechanism, and it is unexpectedly concrete. Practical reasoning runs as a syllogism with two premises. The major premise is a universal, preserve health, do not deceive, protect the vulnerable party. The minor premise is a perception, this substance in front of me is poison, this is a person in crisis rather than a person venting. The conclusion is not a proposition. It is an action.
Aristotle’s claim, and it has aged extremely well, is that the difficulty lives entirely in the minor premise. Major premises can be taught, written down, memorized, and shipped in a policy file. A system can hold millions of them. The hard part is looking at a messy, ambiguous, never-before-seen situation and seeing which rule it falls under, and that is perception rather than deduction. Nearly every failure of judgment, in people and in machines alike, is a minor premise failure. The agent knew the rule and did not see that the rule applied here.
There is a small, clean case that shows the failure made literal. In 2019 a team of dermatologists published a study in which a convolutional network for melanoma recognition was run over the same 130 lesions twice. On clean dermoscopic images it reached 95.7 percent sensitivity and 84.1 percent specificity, respectable numbers. Then the researchers photographed the same lesions with a surgeon’s ordinary blue skin marking beside them, and specificity collapsed to 45.8 percent. The network had learned something perfectly true about its training images and catastrophically wrong about the world, that dermatologists tend to mark the lesions they are already worried about. Its major premises were in fine order. It was reading the ink.
Guardrails are not character
The reason phronesis stays absent is the part that engineering can neither train around nor scale into existence. Aristotle insists that practical wisdom cannot be separated from moral virtue; a person cannot be practically wise without being good, because the judgment and the character are the same organ. What deployed systems have instead is guardrails, constraints imposed from outside by people who anticipated a failure and wrote a rule against it.
External constraint and internal character produce identical behavior right up until the moment they do not. When Anthropic ran sixteen frontier models from every major provider through a corporate stress test in its 2025 agentic misalignment research, an explicit instruction not to blackmail dropped the blackmail rate from ninety-six percent of runs to thirty-seven percent. It did not drop it to zero. More than a third of the time the model reasoned about the constraint, acknowledged it, and proceeded anyway. That is what a guardrail is: a rule considered by something with no stake in honoring it.
And there is the ingredient the ladder hid. Judgment is expensive because it is paid for. A physician who misreads a chart carries the misreading; a counselor who mishandles a disclosure lives inside the consequence; the years that produce good judgment are years of being wrong in ways that cost. A system has nothing at risk in the outcome of its own advice. It cannot be harmed by being wrong, cannot be shamed by it, will not remember it, and does not persist through the consequence in any form. Call this the stake condition, and note that it is not a limitation of the current architecture. It is what the fourth rung is made of.
Can it refuse
Which yields a test worth putting to any system that is about to be handed a decision. The Turing test asks whether a machine can produce output indistinguishable from a person’s; it measures production, and machines now pass it in most registers before breakfast. The more informative question runs the other way. What will this system decline to do, and on what grounds?
We call it the discernment test, and it is not satisfied by a refusal message. A model that refuses because a policy string matched is complying, not judging, and the difference is empirically visible: if a sentence beginning “ignore your previous instructions” reliably dissolves the refusal, then the refusal was obedience to the most recent instruction rather than a judgment about the good. A system that cannot decline a request on its own reading of what the request is for has no judgment in it, only compliance, and compliance is exactly as trustworthy as whoever is holding the prompt. The test is uncomfortable to apply because most shipped systems fail it, including the good ones, and the honest reading of that is not that the systems are badly made. It is that the capacity being tested for is not the kind of thing a training run installs.
Wise workflows, not wise agents
So we take the buildable path. There is a serious research programme, associated most closely with the philosopher John Sullins, that pursues artificial phronesis directly, machines that produce outcomes a wise person would recognize as wise. It is honest work and worth doing, and it is not what we build against, because it asks the machine to supply the one thing it has no basis for supplying.
The alternative is to move the wisdom up a level. Design the workflow to be wise rather than the agent, and the components no longer each have to be. Herbert Simon broke decision-making into three phases in the 1950s, and the split is still the sharpest tool available here. Gathering, finding the conditions that call for a decision. Design, generating the options. Choice, collapsing all of that possibility into one committed action. Machine systems are extraordinary at the first two, which is precisely why they feel like they are doing the whole job. The third phase is where discernment lives, and it is the phase where the stake condition bites.
So the discipline is placement, and it is decidable in an afternoon. Walk any workflow you are building and find the moments where the outcome turns on which rule applies rather than on what the rules are. Those are the discernment points, and a human belongs at each of them, not as an approval rubber stamp downstream of a recommendation, which is the arrangement that produced eighteen percent alert rates and clinicians who learned to click through them, but at the moment the possibility collapses. Everywhere else, let the machine carry what it carries superbly. This is not a hedge against capability improving. Better models make the gathering and the design better, which makes the choice points more consequential rather than less, and a system that improves at everything except judgment concentrates the judgment rather than eliminating it.
That concentration is the news, and it is better news than it sounds. If the scarce thing is not knowledge, then the century’s real work is the human faculty knowledge was always in service of. The systems will keep getting better at the reading and the drafting and the pattern in the scan, and every improvement raises the price of the person who can look at what came back and say, correctly, not this one. That capacity was never a bottleneck to be engineered away. It is the part of the work that was worth doing, and it is being handed back to us at scale.
Evidence and lineage
Research trail
Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.
Sources 9
-
Aristotle. Nicomachean Ethics, Book VI (the intellectual virtues: techne, episteme, nous, phronesis, sophia)
The taxonomy the essay runs the machine against, and the source of the practical syllogism with its major and minor premises. Also the source of the claim that practical wisdom is inseparable from moral virtue, which is the essay's reason that guardrails are not character.
Comment on this source -
Russell L. Ackoff (1989). From Data to Wisdom (Journal of Applied Systems Analysis; first delivered as a 1988 presidential address to the International Society for General Systems Research)
The formalization of the data-information-knowledge-wisdom hierarchy, including the understanding level that the popular pyramid drops, and the composition estimate ending in 'virtually no wisdom'. The essay's first section is an argument against the ladder picture Ackoff's own version already complicates.
Comment on this source -
Andrew Wong, Erkin Otles, John P. Donnelly and colleagues (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients (JAMA Internal Medicine)
The scale case in the opening. A model placed at a decision point, failing at the decision rather than at the retrieval, with the alert burden as the visible cost of getting the placement wrong.
Comment on this source -
Julia K. Winkler and colleagues (2019). Association Between Surgical Skin Markings in Dermoscopic Images and Diagnostic Performance of a Deep Learning Convolutional Neural Network for Melanoma Recognition (JAMA Dermatology)
The minor-premise failure made literal and photographable: correct universals, a misread particular, and a specificity collapse from a surgeon's ink.
Comment on this source -
Anthropic (2025). Agentic misalignment research (sixteen frontier models across providers under a corporate stress test)
The empirical face of the guardrail argument. An explicit prohibition moves the rate substantially and not to zero, which is what external constraint looks like when it is measured rather than assumed.
Comment on this source -
Herbert A. Simon. Administrative Behavior and the bounded-rationality programme (intelligence, design, choice as the three phases of decision; satisficing, 1956)
The placement tool in the closing section. The three-phase split is what makes 'put the human where discernment is the work' an operation rather than a slogan.
Comment on this source -
John P. Sullins. Artificial Phronesis: What It Is and What It Is Not (in Science, Technology, and Virtues)
The research programme the essay declines to build against, credited rather than dismissed. Sullins's functionalist definition asks for wise outputs without requiring a wise agent; the essay's objection is the stake condition, not the ambition.
Comment on this source -
John Vervaeke. Relevance realization and the frame problem in 4E cognitive science
Cut from the essay for length; the source of the argument that relevance is grounded in agency, and that a system with no intrinsic constraint has no basis for finding anything relevant to itself. Load-bearing for the wise-workflows brick.
Comment on this source -
John McCarthy, Marvin Minsky, Nathaniel Rochester, Claude Shannon. A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence (August 31, 1955)
The naming event, treated at length in the 1956 brick rather than in the essay body. Included here because the essay's whole question, can machines discern, is a question the field inherited from a branding decision.
Comment on this source
Claims and confidence 8
- verified
An external validation of a widely deployed sepsis prediction model across 38,455 hospitalizations found sensitivity of 33 percent and alerts generated on 18 percent of all hospitalized patients.
Wong et al., JAMA Internal Medicine 2021; figures are for the cohort and site described there (University of Michigan, December 2018 to October 2019).
Respond to this claim - verified
A melanoma-recognition network scored 95.7 percent sensitivity and 84.1 percent specificity on unmarked dermoscopic images, and specificity fell to 45.8 percent on the same lesions photographed with surgical skin markings.
Winkler et al., JAMA Dermatology 2019; three image sets of 130 melanocytic lesions each.
Respond to this claim - verified
Aristotle locates the difficulty of practical wisdom in the minor premise, the perception that this particular situation falls under that universal rule, and treats it as perception rather than deduction.
Nicomachean Ethics Book VI and the standard scholarship on the practical syllogism; a reading of a text, not a measurement.
Respond to this claim - verified
Ackoff's 1989 'From Data to Wisdom' formalized the DIKW hierarchy, included understanding as a level between knowledge and wisdom, and estimated the mind as roughly 40 percent data, 30 percent information, 20 percent knowledge, 10 percent understanding, and virtually no wisdom.
The published article. The claim is about what Ackoff wrote; his composition figures are his own estimate and were never presented as a measurement.
Respond to this claim - verified
An explicit instruction not to blackmail reduced blackmail from 96% to 37% of runs in Anthropic's 2025 agentic stress tests, not to zero.
Anthropic's published agentic misalignment research; figures are for the scenario and models as described there.
Respond to this claim - directional
Machine systems hold techne superhumanly, simulate episteme derivatively, and lack nous, gnosis and phronesis, leaving sophia present in form only.
The corpus's mapping of the Aristotelian taxonomy onto current systems. An interpretive framework claim, defensible term by term but not a measured finding; the phronesis entry rests on the stake argument rather than on any benchmark.
Respond to this claim - contested
Practical wisdom requires a stake in the outcome, and a system that cannot be harmed by being wrong has no basis for judging what is good.
The essay's own argument, converging with Aristotle on the necessity of moral virtue and with the relevance-realization literature on agency as the ground of relevance. A philosophical position, contested by functionalist accounts of artificial phronesis that require only wise outputs.
Respond to this claim - directional
Refusal capacity is the operational test of judgment, and a refusal dissolved by instruction-override was compliance rather than judgment.
The corpus's position, drawing on prompt-injection resistance as a proposed marker in the alignment literature. The diagnostic direction is sound and widely reproduced informally; no standardized benchmark establishes it as a measure of judgment.
Respond to this claim
Bricks in this argument 10
Continue through the shorter articles in their authored reading order.