⚡
Brief IA
›

AI at Work: Two Clocks and the 400 ms Milestone

🛠️ AI Tools·Tom Levy·

AI at Work: Two Clocks and the 400 ms Milestone

AI at Work: Two Clocks and the 400 ms Milestone
⚡
Key Takeaways
1Under 400 ms, recognition maintains attention and oversight
2Two clocks to manage: 100–400 ms for recognition, the rest for response
3Google targets 200 ms for the next painting; an echo can be sent in 50 ms
💡Why it matters — exceeding the recognition budget causes users to disengage, turns oversight into a hidden cost, and ultimately reshapes the product.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

In AI interfaces, the key metric is no longer just the model's response time. The critical budget hinges on recognition within 400 ms; otherwise, attention wanes and human oversight evaporates. Concrete, quantified, and already tested rules guide the design of these interactions.

When Attention Wanes, Oversight Disappears

In AI tools, the cost of a slow response translates into human supervision rather than just throughput. The recognition budget equates to a budget for controlling outputs: exceeding it means quietly interrupting user verification. In corporate environments, these delays accumulate with the queue of tasks. According to measurements mentioning 117 emails and 153 Teams messages daily, along with interruptions every two minutes, a second lost in recognition can total four minutes of unanswered clicks across 270 instances. In casual chat, users tolerate a slow turn; at work, the queue amplifies the effect. Over time, users return to accept the output rather than verify it, effectively removing the human from the loop. Latency increases the cost of supervision, transforms a collaborative tool into a background task, and ultimately reshapes the product, necessitating dedicated surfaces like notifications and execution history.

What Doherty Measured and Why 400 ms Still Matters

As early as 1982, Walter Doherty and Ahrvind Thadani linked transactions per hour to responsiveness: under approximately 400 milliseconds, productivity rises more than the time saved can explain because the operator maintains their thought instead of reconstructing the context. The threshold says nothing about computational speed; it indicates how long a thought holds. Below 400 ms, pauses do not break the flow; beyond that, errors increase. Other contemporary measurements point in the same direction. Deloitte Digital observed that a gain of 0.1 seconds increased conversions by 8.4% across 37 brands. At the conversational level and across ten languages, the median gap between turns is about 100 ms. On the web, Jakob Nielsen formalized three thresholds: 0.1 s feels instantaneous, 1 s maintains the flow, and 10 s marks the attention limit.

Two Clocks to Manage: Recognition and Final Response

Every AI interaction obeys two clocks. The first starts with the user's action and stops when the interface proves it has heard: echo of the input, button state change, response skeleton. The goal should be 100 ms, with 400 ms as the ceiling. The second clock runs from sending to a usable response; it can last longer if the wait yields a valuable result. What breaks the exchange is the absence of acknowledgment and progress, not the duration. The feedback principles outlined by Ben Shneiderman—real state, permanent cancellation, correction of optimistic actions—remain valid. However, many teams only instrument the second clock driven by model metrics, leaving the first without an owner. This is how a flattering p95 coexists with an inert interface at the first tap. Documenting both objectives and their owners closes the debate on slowness.

Ground Measurements: Google Threshold and Acknowledgments Under 50 ms

Operational benchmarks exist on the interface side. Google considers an interaction-to-next-paint delay of 200 ms or less at the 75th percentile of visits to be "good," regardless of the subsequent inference duration. In most contexts, an interface can issue an acknowledgment within 50 ms, even on modest devices. This acknowledgment can take the form of an echo of the prompt in the thread, a button state change, or a skeleton in the expected response location.

The False Remedy of Streaming and Its Blind Spots

Streaming makes a long response more easily readable in flow, but it does not address the void before the first token. Luke Wroblewski warned as early as 2013 about the effect of the initial delay without feedback. With reasoning or recovery steps, this time can extend while the interface remains unchanged. Worse, displaying a movement or fictitious progress to buy patience backfires on the experience. Users may excuse slowness, but not a misleadingly managed wait.

Agents Add Latency at Every Processing Phase

Agents chain model calls, tools, and prompts, accumulating latency debt at each step. In a contract review scenario—extracting, comparing, reporting, drafting—each phase lasts a few seconds without seeming alarming in isolation. Many interfaces display only a single overall indicator, masking errors that occurred mid-process, which the user only discovers at the end. The documented recommendation is to show the agent's work step by step and allow for stopping. In practice, users initiate execution, switch to Slack, and then return later—much like in the mainframe era. At that point, they validate more than they verify.

Stop Blaming Inference and Set Budgets

Attributing slowness solely to inference time—eight seconds is not uncommon—obscures the interface's responsibility. Generative AI has split the interaction event into recognition and response; the 400 ms threshold applies to the former. Metrics like time to the first token include the model's reflection but cannot substitute for a recognition measure on the surface. Designing on this basis confuses model measurement with interface measurement. The model can take its time to produce the response; the first half-second, however, must prove it has heard.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.