⚡
Brief IA
›

Accounting: AI Advances but Remains Under Supervision

🤖 Models & LLM·Tom Levy·

Accounting: AI Advances but Remains Under Supervision

Accounting: AI Advances but Remains Under Supervision
⚡
Key Takeaways
1AI models are faster, more accurate, and less expensive than accountants for structured tasks
2Claude Opus 5.5, Fable 5.1, and GPT-6 Astra achieve 61.8%, 61.0%, and 57.9% of the criteria, respectively
3Nearly 60% of the benchmark tasks have not been fully resolved, and closing the books requires supervision
💡Why it matters — The study shows rapid advancements in AI in accounting but highlights that essential aspects of the profession remain beyond the reach of current models.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI models outperform accountants on structured tasks, according to Mercor, with the best achieving between 57.9% and 61.8% of the criteria. However, nearly 60% of tasks remain without a complete solution, and book closing still requires supervision. Some key dimensions of the profession have been overlooked.

No Autonomous Closing and a Limited Testing Scope

Mercor indicates that AI models are not capable of closing books without supervision. No model has fully solved nearly 60% of the benchmark tasks. The tested tasks correspond to what AI handles best, such as accurately following instructions and searching for details. Essential aspects of the profession, such as communication with clients, collaboration with colleagues, or utilizing context accumulated over the years, were not evaluated. According to Mercor, these omissions explain why accountants cannot be replaced, even though significant productivity gains are expected in the sector.

Current Scores and Leap in Eighteen Months

Claude Opus 5.5 achieves the highest score with 61.8% of the rating criteria met, followed by Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%. Eighteen months ago, the top-performing AI models had results below the average for accountants, which was set at around 37%. Now, these same models are able to handle simplified tasks almost perfectly.

What APEX Measures and How the Comparison Was Made

The APEX benchmark consists of 160 exercises spread across 10 fictitious companies, designed by over 40 specialists with an average of 11 years of experience. As part of the study, twelve certified public accountants, averaging five and a half years of experience, performed simplified versions of these tasks. For these standardized exercises, Mercor indicates that AI systems outperform professionals in speed, accuracy, and cost. The accuracy of the AIs was compared to that of accountants operating without assistance, with detailed results according to the launch date of the models.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.