Brief IA

QIMMA: A Revolution in the Evaluation of Arabic LLMs

🔬 Research·Tom Levy·

QIMMA: A Revolution in the Evaluation of Arabic LLMs

QIMMA: A Revolution in the Evaluation of Arabic LLMs
Key Takeaways
1QIMMA introduces rigorous benchmark validation to evaluate Arabic LLMs, ensuring genuine linguistic capability.
2Over 52,000 samples from 109 subsets are used to test the models across 7 diverse domains.
3QIMMA's methodology reveals recurring quality issues in the benchmarks, including factual errors and cultural biases.
💡Why it mattersQIMMA enhances the reliability of evaluations for Arabic LLMs, crucial for accurate and culturally sensitive applications.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

QIMMA قِمّة: An Arab LLM Ranking Focused on Quality

QIMMA قِمّة stands out for its ability to validate benchmarks before evaluating models, ensuring that the reported scores reflect the true linguistic competence in Arabic of the language models.

The Problem: Fragmentation and Lack of Validation

Arabic is a language spoken by over 400 million people, encompassing a multitude of dialects and cultural contexts. However, the evaluation of natural language processing (NLP) for Arabic is currently fragmented. Several major issues have motivated the development of QIMMA:

  • Translation Issues: Many Arabic benchmarks are simply translations from English, leading to distributional mismatches. Questions that seem natural in English often become awkward or culturally inappropriate when translated into Arabic.

  • Lack of Quality Validation: Even benchmarks created directly in Arabic are often published without undergoing rigorous quality checks. This results in inconsistencies in annotations, incorrect answers, encoding errors, and cultural biases in ground truth labels.

  • Reproducibility Gaps: Evaluation scripts and sample results are rarely published, making it difficult to audit results or build upon previous work.

  • Coverage Fragmentation: Existing rankings often focus on isolated tasks and narrow domains, making comprehensive evaluation of models challenging.

What QIMMA Contains

QIMMA aggregates 109 subsets from 14 source benchmarks into a unified evaluation suite of over 52,000 samples, covering seven domains:

  • Culture: AraDiCE-Culture, ArabCulture, PalmXMCQ
  • STEM: ArabicMMLU, GAT, 3LM STEMMCQ
  • Legal: ArabLegalQA, MizanQAMCQ, QA
  • Medical: MedArabiQ, MedAraBenchMCQ, QA
  • Security: AraTrustMCQ
  • Poetry & Literature: FannOrFlopQA
  • Coding: 3LM HumanEval+, 3LM MBPP+Code

The Quality Validation Pipeline

At the heart of QIMMA is a multi-step quality validation pipeline. Before running any model, each sample from each benchmark undergoes this rigorous validation.

  • Step 1: Multi-Model Automated Evaluation
    Each sample is independently evaluated by two state-of-the-art language models: Qwen3-235B-A22B-Instruct and DeepSeek-V3-671B. Each model assigns a score to a sample based on a 10-point rating scale.

  • Step 2: Human Annotation and Review
    Flagged samples are then reviewed by native Arabic speakers with cultural and dialectal familiarity. Human annotators make the final decisions on cultural context and regional variation.

What We Found: Systematic Quality Issues

The pipeline revealed recurring quality issues across the benchmarks, reflecting gaps in how the benchmarks were initially constructed.

  • Response Quality: False or poorly matched gold indices, factually incorrect answers, missing or raw responses.
  • Text and Formatting Quality: Corrupted or unreadable text, spelling and grammar errors, duplicated samples.
  • Cultural Sensitivity: Reinforcement of stereotypes and monolithic generalizations about diverse communities.

Ranking Results

The results, published in April 2026, cover the top 10 evaluated models. The Qwen/Qwen3.5-397B-A17B-FP8 model ranks first with a score of 68.06. Arabic-specialized models excel in cultural and linguistic tasks, while coding remains the most challenging domain for these models. The size of a model does not necessarily guarantee its performance, highlighting the importance of linguistic and cultural specialization.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.