Skip to content

Incentives · research february to july 2026 · published 2026-08-03 · v1 · 3 min read

Silence is an intervention

A product that refuses to answer instantly is protecting something speed destroys

Why we decline to optimize time to first token in relational contexts, from the response-time law that governs the rest of software to the delays our specifications require on purpose.

In brief
The problem

verified

Every claim this passage rests on has been checked against its sources.

  • "IBM researchers Walter Doherty and Arvind Thadani published a 400 millisecond response-time threshold in 1982, and it entered interface practice as a standard target."

    verified. The 1982 IBM Systems Journal paper and four decades of its citation in interaction design practice.

Open the complete evidence in the structured publication.

Four decades of interface practice treat every millisecond of delay as pure loss, and the instinct arrives unexamined in contexts where an instant reply says the wrong thing.
The mechanism

verified

Every claim this passage rests on has been checked against its sources.

  • "Our design system specifies that the system does not optimize for minimum time to first token in session contexts, and requires a presence indicator to hold for a minimum duration before the first text renders."

    verified. Our interaction design system document, read from the record before authoring. A statement about what our documents specify, not a measured outcome.

Open the complete evidence in the structured publication.

The interval before a reply is itself a message about the kind of attention being paid, so a system that drives the interval to zero has not removed a cost, it has overwritten what the pause was saying.
The move

position

This is the publication's stated position, not an empirical claim. It rests on the argument rather than graded evidence.

Open the complete evidence in the structured publication.

Treat latency as a budget rather than a defect, and spend it deliberately wherever a person is being met rather than served.

The orthodoxy has receipts, and they are good ones. In 1982 two IBM researchers, Walter Doherty and Arvind Thadani, published a 400 millisecond response-time threshold, the point below which a person stops waiting on a machine and starts thinking with it, and the number still sets targets in interface work forty years later. Google measured the same law from the other side in 2009, when a controlled experiment adding between a tenth and four tenths of a second of delay to search results reduced the number of searches people ran, with the effect persisting for weeks after the delay was removed. Latency is a tax. We are not disputing the finding.

We are disputing where it generalizes. Doherty measured programmers editing code. The Google experiment measured people looking things up. Both are transactions, tasks with a right answer and a person who wants it and then wants to be gone, and in a transaction the wait is pure overhead, which is exactly why removing it is pure gain. Carry the same instinct into a moment where somebody has just said something difficult, and the technically excellent instant reply communicates that nothing happened in between. The interval before a reply is itself a message about the kind of attention being paid, which means a system that drives the interval to zero has not removed a cost, it has overwritten what the pause was saying.

So the delay is written into the specification rather than left to taste, because a number nobody wrote down loses every argument with a performance dashboard. Our design system states outright that the system does not optimize for minimum time to first token in session contexts, and requires a presence indicator to hold for a minimum duration before the first character of text appears, so that an immediate model response does not arrive with the uncanny quality of an answer that was waiting. The motion specification sets a default initial response delay of roughly a second and a half, extended to between two and four seconds in high-arousal emotional contexts, and paces streamed text at around three words per second with pauses at sentence and clause boundaries rather than revealing it token by token. Before any waiting indicator appears at all there is an attentive silence window of three to five seconds, and a companion rule that silence is never left past five to seven seconds without a minimal acknowledgment, because unmarked absence stops reading as presence and starts reading as failure.

The doses are the argument. Slowness is not a virtue here, and an interface that made a person wait to seem thoughtful would be running the same manipulation as the one that answers instantly to seem capable, only with worse manners. That is why the same specification bypasses the silence behavior entirely for urgent or safety-critical content, and why the pacing targets vary by phase rather than sitting at one pious constant, faster for structured task work and slower where something is being processed. The pause is instrumentation for a particular kind of moment, not a house style.

Every one of those numbers is a number a dashboard would flag as a regression. Time to first token is the metric the whole industry ships against, and the honest description of what we are doing is that we accept a worse score on it in the contexts where the score measures the wrong thing. This is what designing against your own incentives looks like in practice, not a manifesto but a default in a spec file that costs something real on a chart somebody reports upward every month. A product cannot protect a pause it has not agreed, in writing and in advance, to lose an argument about, and the pauses worth protecting are the ones where a person is deciding what they actually think.

Evidence and lineage

Research trail

Follow the sources, inspect how the claims are graded, or propose a correction at the exact record it concerns.

Sources 3
  1. Walter J. Doherty and Arvind J. Thadani (1982). The Economic Value of Rapid Response Time (IBM)

    The received view at its strongest and best evidenced. The 400 millisecond figure is the origin of the industry's latency doctrine, and the brick concedes it in full before restricting its domain.

    Comment on this source
  2. Jake Brutlag (Google) (2009). Speed Matters for Google Web Search

    The controlled measurement of the other side of the same law, and the source of the persistence finding. Named because the brick's argument depends on the experiments being sound rather than on them being wrong.

    Comment on this source
  3. MNSTRY design system and platform documentation (2026). Interaction design system, experience design specification, and attentive motion and loading specification (internal records of the pacing contract)

    The shipped counter-position. The stated refusal to optimize time to first token in session contexts, the minimum presence duration, the initial response delays, the streaming pace, the silence window, and the safety carve-out.

    Comment on this source
Claims and confidence 6
  1. verified

    IBM researchers Walter Doherty and Arvind Thadani published a 400 millisecond response-time threshold in 1982, and it entered interface practice as a standard target.

    The 1982 IBM Systems Journal paper and four decades of its citation in interaction design practice.

    Respond to this claim
  2. verified

    A 2009 Google experiment adding between 100 and 400 milliseconds of delay to search results reduced the number of searches users performed, with the effect persisting for weeks after the delay was removed.

    Google's own published account of the experiment, presented at Velocity 2009.

    Respond to this claim
  3. verified

    Our design system specifies that the system does not optimize for minimum time to first token in session contexts, and requires a presence indicator to hold for a minimum duration before the first text renders.

    Our interaction design system document, read from the record before authoring. A statement about what our documents specify, not a measured outcome.

    Respond to this claim
  4. verified

    Our motion specification sets a default initial response delay of about 1500 milliseconds, extends it to between 2000 and 4000 milliseconds in high-arousal emotional contexts, and paces streamed text at roughly three words per second with pauses at sentence and clause boundaries.

    The same internal specification, which states these as canonical tokens and phase targets.

    Respond to this claim
  5. verified

    Our motion specification names an attentive silence window of three to five seconds before any waiting indicator appears, and forbids leaving silence beyond five to seven seconds without a minimal acknowledgment.

    The attentive silence policy in the same specification.

    Respond to this claim
  6. verified

    The same specification bypasses the silence behavior for urgent or safety-critical user content.

    Stated as an explicit exception in the attentive silence policy.

    Respond to this claim

Read next

Or survey the topics.

Concepts in this piece 1

Add to the work

Contribute to Silence is an intervention

Write the useful part. Identity, provenance, and review history are attached when you submit. The published source stays unchanged.

Target Silence is an intervention

Contribution intent
Use an agent instead

The interface is ready. Public authenticated intake remains off until the hosted migration and feature flag are deployed together.