QIMMA: A Revolution in the Evaluation of Arabic LLMs
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
QIMMA قِمّة: An Arab LLM Ranking Focused on Quality
QIMMA قِمّة stands out for its ability to validate benchmarks before evaluating models, ensuring that the reported scores reflect the true linguistic competence in Arabic of the language models.
The Problem: Fragmentation and Lack of Validation
Arabic is a language spoken by over 400 million people, encompassing a multitude of dialects and cultural contexts. However, the evaluation of natural language processing (NLP) for Arabic is currently fragmented. Several major issues have motivated the development of QIMMA:
-
Translation Issues: Many Arabic benchmarks are simply translations from English, leading to distributional mismatches. Questions that seem natural in English often become awkward or culturally inappropriate when translated into Arabic.
-
Lack of Quality Validation: Even benchmarks created directly in Arabic are often published without undergoing rigorous quality checks. This results in inconsistencies in annotations, incorrect answers, encoding errors, and cultural biases in ground truth labels.
-
Reproducibility Gaps: Evaluation scripts and sample results are rarely published, making it difficult to audit results or build upon previous work.
-
Coverage Fragmentation: Existing rankings often focus on isolated tasks and narrow domains, making comprehensive evaluation of models challenging.
What QIMMA Contains
QIMMA aggregates 109 subsets from 14 source benchmarks into a unified evaluation suite of over 52,000 samples, covering seven domains:
- Culture: AraDiCE-Culture, ArabCulture, PalmXMCQ
- STEM: ArabicMMLU, GAT, 3LM STEMMCQ
- Legal: ArabLegalQA, MizanQAMCQ, QA
- Medical: MedArabiQ, MedAraBenchMCQ, QA
- Security: AraTrustMCQ
- Poetry & Literature: FannOrFlopQA
- Coding: 3LM HumanEval+, 3LM MBPP+Code
The Quality Validation Pipeline
At the heart of QIMMA is a multi-step quality validation pipeline. Before running any model, each sample from each benchmark undergoes this rigorous validation.
-
Step 1: Multi-Model Automated Evaluation
Each sample is independently evaluated by two state-of-the-art language models: Qwen3-235B-A22B-Instruct and DeepSeek-V3-671B. Each model assigns a score to a sample based on a 10-point rating scale. -
Step 2: Human Annotation and Review
Flagged samples are then reviewed by native Arabic speakers with cultural and dialectal familiarity. Human annotators make the final decisions on cultural context and regional variation.
What We Found: Systematic Quality Issues
The pipeline revealed recurring quality issues across the benchmarks, reflecting gaps in how the benchmarks were initially constructed.
- Response Quality: False or poorly matched gold indices, factually incorrect answers, missing or raw responses.
- Text and Formatting Quality: Corrupted or unreadable text, spelling and grammar errors, duplicated samples.
- Cultural Sensitivity: Reinforcement of stereotypes and monolithic generalizations about diverse communities.
Ranking Results
The results, published in April 2026, cover the top 10 evaluated models. The Qwen/Qwen3.5-397B-A17B-FP8 model ranks first with a score of 68.06. Arabic-specialized models excel in cultural and linguistic tasks, while coding remains the most challenging domain for these models. The size of a model does not necessarily guarantee its performance, highlighting the importance of linguistic and cultural specialization.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.