⚡
Brief IA
›

LLM: Small Windows, Tight Budget, and Sliding Window

🤖 Models & LLM·Tom Levy·

LLM: Small Windows, Tight Budget, and Sliding Window

LLM: Small Windows, Tight Budget, and Sliding Window
⚡
Key Takeaways
1Allocate 20% to the system, 20% to history, and 60% to recovery, capping the prompt and stopping additions once the limit is reached.
2A sliding window removes the oldest turns, fixes the processed size, and stabilizes latency.
3A heuristic suggests 1 word = 1.3 tokens, and an example shows a cutoff at 30 words.
💡Why it matters — These methods prevent large documents from overwhelming the prompt and ensure predictable usage of tokens and latency.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

In the practical uses of language models, streamlined contextual approaches aim to reduce costs and latency without overwhelming the prompt. Two key levers emerge: a strict token budget allocation, coupled with document retrieval, and the controlled removal of the oldest turns. Python examples illustrate these mechanisms and their parameters.

Allocate 20% to the system, 20% to history, 60% to data

One method involves dividing the context window into zones with explicit ceilings. One example allocates up to 20% for system instructions, 20% for chat history including the latest request, and the remaining 60% for retrieved content. This division mandates the cessation of insertion when a zone reaches its limit, preventing a massive document from consuming the prompt and helping to mitigate the "lost in the middle" effect. A Python example demonstrates the construction of a prompt under word constraints: it calculates a base cost for the system message and request, iterates through the retrieved pieces, adds them only if the ceiling is not exceeded, and then stops as soon as the budget is reached. The selected elements are assembled into a single context before being combined with the system message and the request. A test sets a very small budget of 30 words to enforce truncation, aiming to include the phrase "Madrid is the capital of Spain" among documents that also mention Seville, the geographical situation of the country, and a population of around 47 million. This logic fits within augmented generation through retrieval, which enriches the prompt with relevant external documents.

Removing old turns stabilizes latency and token usage

The sliding window treats history as a FIFO queue: with each new message, the oldest ones are removed. The window size parameter serves as a slider between context retention and load management, resulting in limited but predictable memory. The proposed benefit lies in the strict control of tokens and computational load, with the number of exchanges considered remaining fixed, which stabilizes latency. In the Python example, a class maintains a bounded history with a max_turns initialized to 3 by default. Upon addition, if the length exceeds the limit, the history is truncated to the last turns. The prompt construction includes a system reminder about conciseness based on recent context and recomposes the user–AI pairs still present. A test with max_turns=2 shows that only the most recent interactions are taken into account, reducing the prompt to the bare minimum.

Limiting the window targets costs, latency, and contextual noise

In practical deployments, massively extending the context window exposes users to high API costs and response times deemed unacceptable. Additionally, there is a risk that the model may overlook buried elements within a long prompt. Conversely, restricting the window, if managed with clear rules, can reduce latency and costs while helping the model focus on information relevant to the response. These expected gains rely on explicit management choices: zone limits, cessation of insertion upon exceeding limits, or controlled removal of too-old turns. They aim to achieve the same objectives: avoid exhausting the context budget and contain noise that dilutes relevance.

Useful parameters and heuristics for implementation

The examples suggest concrete adjustments. For budgeting, the prompt construction function takes a word ceiling, with a default illustrative value of 50, and relies on a simple word count as a proxy for tokens. To improve alignment with actual costs, a heuristic proposes that one word corresponds, on average, to 1.3 tokens. On the sliding window side, the example class sets max_turns to 3 by default and demonstrates the effect of a tighter threshold with max_turns=2. The set of simulated interactions ranges from discovering Python to questions about lists and their types, then proves that only the most recent remain in the prompt, in accordance with the imposed limit.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.