⚡
Brief IA
›

A Protocol Assesses the Reliability of LLMs Without Ground Truth

🔬 Research·Tom Levy·

A Protocol Assesses the Reliability of LLMs Without Ground Truth

A Protocol Assesses the Reliability of LLMs Without Ground Truth
⚡
Key Takeaways
1Consistency between successive executions of an AI agent serves as a proxy to estimate reliability in the absence of ground truth
2The command-line tool cca allows for comparing the consistency of different models and agents on coding tasks
3Sources of non-determinism include sampling, temperature, system effects, and the lack of controls in certain tools
💡Why it matters — Multi-execution consistency provides an operational criterion for assessing the reliability of AI agents in contexts where the correct answer is not known in advance.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

In coding and analysis practices, the responses of language models vary from one execution to another. A protocol based on the consistency observed across multiple runs provides a reliability benchmark when no ground truth is available. A command-line tool, cca, allows for the comparison of models and agents on coding problems and explores this proxy further.

Comparing multiple responses helps estimate model reliability

When a ground truth is lacking, the consistency between successive executions can serve as an indicator of reliability. For a given class of problems, if a tool produces different solutions that consistently evaluate to the same correct result, this pattern can be reused as a benchmark for a new problem of the same class. Observing distinct solutions converge again towards the same answer reinforces the trust placed in that answer. The framework presented specifically aims to exploit this regularity by treating multi-sample consistency as a proxy for reliability.

Where the output divergences of LLMs come from

Language models generate sequences token by token in an autoregressive mode. Self-attention assembles the context, propagated through feed-forward layers powered by billions of weights, to produce logits per token, which are converted into probabilities via softmax. In practice, sampling randomly draws a token according to this distribution, potentially truncated by top-k or top-p, so that a slight deviation at the beginning of the sequence can significantly influence the rest. The temperature, applied to the logits, modulates this distribution: close to 0, the probability of the best token approaches 100%, resembling greedy sampling; beyond 1, the distribution flattens, variability increases, and consistency decreases. Even with zero temperature, effects like numerical rounding, batching, and parallelism can introduce residual variations.

Sampling controls sometimes inaccessible

Several reasoning models, such as GPT5+ and Claude Opus, do not necessarily expose temperature or sampling settings. Coding tools like Claude Code, Codex, or Antigravity also do not make them available to developers. Reasoning models require finely tuned non-zero settings to function well, and these parameters can be adjusted on the fly depending on the question or generated content. As model power increases, understanding and influencing non-determinism becomes more challenging.

A CLI to test and compare agent consistency

The command-line tool Coding Agent Consistency (cca) orchestrates repeated experiments and analyzes the distribution of results. The demonstrations rely on well-defined coding problems with smaller models to contain costs, but the approach can be extended to other uses. The package allows for exploring models via API with litellm, interacting with local models via ollama, and querying coding agents via Omnigent. It thus becomes possible to compare the consistency of multiple tools on the same problem or across a set of problems.

When the industry adopts agents: reliability requirements

The rise of coding agents and AI analysts accelerated between 2025 and early 2026, with the adoption of tools like Claude Code, Codex, and Antigravity, capable of chaining multiple model calls per request to solve complex tasks. In business analysis, assistants sift through datasets, write SQL queries, and produce numerical and recommended reports. This shift, which frees users from implementation details to focus on design, necessitates the ability to measure reliability. Model-problem combinations exhibit different consistency profiles, highlighting the need for comparative tools and protocols.

Business use cases, auditability, and practical limits

In single-answer tasks, measuring consistency is crucial, especially since deployments rarely have a ground truth to verify results. In business analysis, framing a request like the distribution of lifetime value by age calls for explicit business definitions. A trustworthy system should produce identical figures for the same question, even if the phrasing varies, a challenging goal to achieve with LLMs. In contrast, a human analyst documents their procedure, which can become a team reference after review. Querying an agent ten times is akin to questioning ten independent or semi-independent analysts based on history and memory, especially if the temperature is non-zero and the question is vague. Finally, a wealth of literature already addresses metrics of consistency and solution diversity, while non-determinism, a source of adaptability and creativity, necessitates reasoning in terms of consistency defined as the reproducibility of an output for the same input, in systems where long chains of reasoning and unseen parameters complicate analysis.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.