Constrained Decoding Provides Unprecedented Formal Guarantees

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Decoding constraints provide formal guarantees that prompts do not, but they can alter a model's distribution and harm quality in strict situations. Their implementation relies on token masks, which are difficult and sometimes costly to construct, with ongoing work on the GPU side. Pragmatic usage schemes are proposed to leverage them without degrading reasoning.
Under constraint, validity does not always rhyme with quality
An output that conforms to a scheme does not guarantee its probability or truthfulness. The application of greedy local masking can modify the model's distribution and lead to a decrease in output quality or reasoning performance when constraints are strict. Studies have compared different constraint methods and documented their effects. Solutions like ASAP are proposed to limit the negative impact of masking. Some metrics can become misleading, as they are rendered trivially favorable by the application of constraints.
What constraint imposes that the prompt does not ensure
Prompting affects probabilities through context but cannot make the selection of invalid tokens impossible. In contrast, constrained decoding applies a mask that excludes choices that violate a formal rule, such as grammar or a JSON schema. In practice, asking the model to produce only valid JSON with examples in the prompt works in 99 out of 100 calls, but this does not constitute a formal guarantee at the token level.
Implementing masks and adopting pragmatic usages
The main difficulty of constrained decoding lies in constructing the mask of allowed tokens, a step that can be costly. Research is focused on accelerating this preprocessing and generating masks suitable for GPUs. For practical use, it is advisable to let the model reason freely and then constrain the extraction, adjusting the rigor of the constraints according to the task. Recommendations specify in which cases constrained decoding is relevant and how to interpret its guarantees.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.