AI at Work: Two Clocks and the 400 ms Milestone

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
In AI interfaces, the key metric is no longer just the model's response time. The critical budget hinges on recognition within 400 ms; otherwise, attention wanes and human oversight evaporates. Concrete, quantified, and already tested rules guide the design of these interactions.
When Attention Wanes, Oversight Disappears
In AI tools, the cost of a slow response translates into human supervision rather than just throughput. The recognition budget equates to a budget for controlling outputs: exceeding it means quietly interrupting user verification. In corporate environments, these delays accumulate with the queue of tasks. According to measurements mentioning 117 emails and 153 Teams messages daily, along with interruptions every two minutes, a second lost in recognition can total four minutes of unanswered clicks across 270 instances. In casual chat, users tolerate a slow turn; at work, the queue amplifies the effect. Over time, users return to accept the output rather than verify it, effectively removing the human from the loop. Latency increases the cost of supervision, transforms a collaborative tool into a background task, and ultimately reshapes the product, necessitating dedicated surfaces like notifications and execution history.
What Doherty Measured and Why 400 ms Still Matters
As early as 1982, Walter Doherty and Ahrvind Thadani linked transactions per hour to responsiveness: under approximately 400 milliseconds, productivity rises more than the time saved can explain because the operator maintains their thought instead of reconstructing the context. The threshold says nothing about computational speed; it indicates how long a thought holds. Below 400 ms, pauses do not break the flow; beyond that, errors increase. Other contemporary measurements point in the same direction. Deloitte Digital observed that a gain of 0.1 seconds increased conversions by 8.4% across 37 brands. At the conversational level and across ten languages, the median gap between turns is about 100 ms. On the web, Jakob Nielsen formalized three thresholds: 0.1 s feels instantaneous, 1 s maintains the flow, and 10 s marks the attention limit.
Two Clocks to Manage: Recognition and Final Response
Every AI interaction obeys two clocks. The first starts with the user's action and stops when the interface proves it has heard: echo of the input, button state change, response skeleton. The goal should be 100 ms, with 400 ms as the ceiling. The second clock runs from sending to a usable response; it can last longer if the wait yields a valuable result. What breaks the exchange is the absence of acknowledgment and progress, not the duration. The feedback principles outlined by Ben Shneiderman—real state, permanent cancellation, correction of optimistic actions—remain valid. However, many teams only instrument the second clock driven by model metrics, leaving the first without an owner. This is how a flattering p95 coexists with an inert interface at the first tap. Documenting both objectives and their owners closes the debate on slowness.
Ground Measurements: Google Threshold and Acknowledgments Under 50 ms
Operational benchmarks exist on the interface side. Google considers an interaction-to-next-paint delay of 200 ms or less at the 75th percentile of visits to be "good," regardless of the subsequent inference duration. In most contexts, an interface can issue an acknowledgment within 50 ms, even on modest devices. This acknowledgment can take the form of an echo of the prompt in the thread, a button state change, or a skeleton in the expected response location.
The False Remedy of Streaming and Its Blind Spots
Streaming makes a long response more easily readable in flow, but it does not address the void before the first token. Luke Wroblewski warned as early as 2013 about the effect of the initial delay without feedback. With reasoning or recovery steps, this time can extend while the interface remains unchanged. Worse, displaying a movement or fictitious progress to buy patience backfires on the experience. Users may excuse slowness, but not a misleadingly managed wait.
Agents Add Latency at Every Processing Phase
Agents chain model calls, tools, and prompts, accumulating latency debt at each step. In a contract review scenario—extracting, comparing, reporting, drafting—each phase lasts a few seconds without seeming alarming in isolation. Many interfaces display only a single overall indicator, masking errors that occurred mid-process, which the user only discovers at the end. The documented recommendation is to show the agent's work step by step and allow for stopping. In practice, users initiate execution, switch to Slack, and then return later—much like in the mainframe era. At that point, they validate more than they verify.
Stop Blaming Inference and Set Budgets
Attributing slowness solely to inference time—eight seconds is not uncommon—obscures the interface's responsibility. Generative AI has split the interaction event into recognition and response; the 400 ms threshold applies to the former. Metrics like time to the first token include the model's reflection but cannot substitute for a recognition measure on the surface. Designing on this basis confuses model measurement with interface measurement. The model can take its time to produce the response; the first half-second, however, must prove it has heard.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.