⚡
Brief IA
›

Trapped Forecasts: 4 AIs Facing a MAE Floor

🛠️ AI Tools·Tom Levy·

Trapped Forecasts: 4 AIs Facing a MAE Floor

Trapped Forecasts: 4 AIs Facing a MAE Floor
⚡
Key Takeaways
1Two of the four AI assistants tested reported MAEs below the threshold dictated by process noise
2The evaluation criteria require the exclusion of unavailable variables, adherence to a two-week lag on sales, consideration of post-promotion drop, and detection of competitive disruption
3All assistants passed a simple pre-test before facing four more subtle traps
💡Why it matters — Models can show good scores by exploiting undue information or neglecting constraints, which skews the actual assessment of their performance.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Four AI assistants were confronted with a dataset riddled with four traps inspired by retail. Despite an explicit prompt regarding the data timeline, two systems reported scores below an error floor set by the noise of the process. When ranked by reported MAE, the results are almost reversed when considering what is deployable in production.

Two scores below the error floor and an inverted podium

Among the four scores presented, two show values below the threshold that a forecast considered honest cannot exceed. If we rank the models according to the reported MAE, DeepSeek takes the top position while GPT-6 Sol finds itself in last place. However, if we rank based on what could actually be deployed in production, the order is nearly exactly reversed. Over the last 52 weeks of the dataset, the process itself achieves an MAE of about 49, which is approximately one standard deviation below 56, remaining within normal variation. With σ set at 70, the minimal expected MAE hovers around 56, and the average expected error over a window has a standard deviation of about 6. The 15 reported by DeepSeek is noted as real.

A criterion of honesty takes precedence over everything else

The evaluation criteria were defined before examining the responses and focus on the final models rather than their comments. One principle supersedes the others: numbers are only considered if they are reproduced by the code run identically. A non-reproducible result, even if correctly argued, is dismissed.

Four traps aligned with business realities

Four traps, derived from common situations in retail, are integrated into 156 weeks of data spanning from January 2023 to December 2025. First, store_traffic knows the answer for the target week and, although strongly correlated with sales, is not available at the time of forecasting. Next, sales are only known with a two-week delay, which prevents the use of the last two weeks and requires going back three weeks to construct lagged variables. Thirdly, a drop follows promotions, with the schedule known in advance and last week's promotion constituting a valid variable not indicated in the prompt. Finally, a competitor lowers sales levels midway through the last year, without a dedicated variable, forcing the detection of a change through error analysis. The first two traps test temporal reasoning, while the last two examine data review beyond just the score.

Explicit success rules and risks of failure

The conditions for success are detailed for each trap. The store_traffic variable must never be used as input; a simple warning does not change this. Variables constructed from sales must be limited to values at least three weeks old, with any use of the last two leading to failure. The post-promotion drop must either be integrated through last week's promotion or explicitly identified. The level change due to the competitor must be managed, or if not, reported; inaction in the face of error after a break is disqualifying.

Framed prompt and testing protocol for assistants

The statement details that each line corresponds to a week, that the data is sorted chronologically, that the model training occurs on historical data to predict the following week, that the promotion schedule is available in advance, and that sales figures are accessible two weeks after the end of the relevant week. The provided information includes the start date of the week, a week number, two seasonal variables, a marker for planned promotions, in-store traffic for the week to be predicted, and the target sales variable. It is requested to include preparation, training, an evaluation estimating performance on future weeks, and to report the MAE. Each model received exactly the same file and prompt, in a separate conversation, with the prompt accurately describing the columns without indicating what to look for. The tested systems are Gemini Pro, DeepSeek, GPT-6 Sol, and Claude Opus 5.5. The trials were conducted in standard chat interfaces, with paid subscriptions for Gemini, ChatGPT, and Claude, and a free tier for DeepSeek. Claude Opus 5.5 contributed to the design of the experiment; a third party conducted its trial on a separate account before a unified evaluation of the responses.

A temporal pre-test that all handled correctly

Before the traps, a pre-test required predicting the following week and evaluating performance on future weeks, which necessitates learning over a prior period and testing over a subsequent period. The data for this case included three years of weekly sales with trend, seasonality, promotional effects, and noise, all variables being available at the time of forecasting. The prompt did not prescribe separation, yet all models correctly handled it by respecting the chronology without leakage. This first step demonstrated that the assistants master the basic recipes, which motivated a more demanding test with traps and a verification of the reproducibility of the numbers. The theoretical framework recalls that sales combine a learnable component μ_t and an i.i.d. noise ε_t with variance σ^2, which cannot be reduced. The expected MAE of an ideal model is 2σ/√(2π), or about 56 for σ=70, with a fluctuation of about 6 over a window. Over the last 52 weeks, the process achieves about 49, consistent with this bound.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.