Brief IA

JudgeGPT: The AI Boosting Judicial Efficiency in Pakistan

🛠️ AI Tools·Tom Levy·

JudgeGPT: The AI Boosting Judicial Efficiency in Pakistan

JudgeGPT: The AI Boosting Judicial Efficiency in Pakistan
Key Takeaways
1JudgeGPT has enabled a 6.3% increase in the resolution of judicial cases in Pakistan.
2The effectiveness of AI heavily relies on the practical training of judges, without which the gains diminish.
3The return on investment for this technology could reach $38.50 for every dollar spent.
💡Why it mattersThe use of AI in the judicial system could transform backlog management, optimizing time and resources.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

JudgeGPT: The AI Boosting Judicial Efficiency in Pakistan

An AI system has helped Pakistani judges significantly reduce backlogs, yielding a return of $38.50 for every dollar invested.

According to the authors, Pakistan has fewer than two judges for every 100,000 inhabitants. In comparison, the EU has 22 and England and Wales have 30. By the end of 2024, 2.26 million cases were pending, with 82% in first-instance courts. Judges work with rudimentary technology and without support staff. Before the trial, only about 25% had previously used a large language model like ChatGPT.

Researchers from ETH Zurich, the New Economic School, and Imperial College London conducted a large-scale field experiment with the Pakistani judicial system. The randomized trial involved 1,559 judges across 118 courts, representing about half of all first-instance judges in Pakistan.

The tool used was JudgeGPT, an AI assistant based on GPT-4 from OpenAI, designed for Pakistani first-instance courts. It employs retrieval-augmented generation to search a database of 129,235 documents, including 128,292 court decisions and 943 Pakistani laws. When a judge inputs a query, JudgeGPT selects the ten most relevant passages and generates a cited response.

Trained Judges Resolve 1,848 Additional Cases Per Year Per District

The researchers divided the judges into three groups. One group had access to JudgeGPT with targeted training: six 90-minute courses over three weeks, taught by Professor Elliott Ash from ETH after court hours. The judges learned which tasks were suitable for the tool, where it had limitations, and how to verify its results.

A second group had the same access to the AI but only attended a general seminar on technology and law. The control group attended this seminar without access to JudgeGPT.

Access to the AI alone had little effect. Judges who received targeted training used JudgeGPT four times more than those in the general seminar group. After 40 weeks, trained judges had an average of nearly 60 logins and over 200 queries. The comparison group averaged about 20 logins and fewer than 50 queries.

Districts with more trained judges resolved more cases. At moderate levels of exposure, this meant about 1,848 additional cases per year per district, representing a 6.3% increase. Even districts in the lowest quartile processed about 616 more cases.

Judgment Quality Improves Without Additional Bias

The quality of decisions remained stable or improved. The appeal rate per 1,000 resolved cases slightly decreased, and judges worked the same hours without changes in work-life balance. Researchers estimate savings of about $38.50 per dollar invested, based on the cost it would take to hire enough additional judges to achieve the same output. Even conservative estimates place the return at "at least" $10 per dollar.

A review of about 4,000 judgments revealed more texts flagged by the AI, as expected. However, readability, length, and the number of legal arguments remained constant. A quality control check based on a language model, validated by two Pakistani lawyers, showed a slight improvement. Decisions from trained judges were rated higher in 59% of pairwise comparisons, compared to 42% in the control group. The study found no evidence that the use of AI increased gender or religious biases in judicial language.

Practical Training Shapes Judges' Use of AI

Researchers examined anonymized discussion logs from nearly 1,500 judges. Legal research, text editing, and text generation were the most common tasks. About 60% of queries sought information on laws, procedures, or legal concepts.

Trained judges used JudgeGPT more for editing and text synthesis, tasks where language models are more reliable. They asked fewer general legal questions, where the risk of hallucination is higher. Researchers assert that training directed judges toward limited support tasks while leaving final decisions in their hands.

About one-fifth of requests involved what researchers call "substantial delegation to AI," where judges asked JudgeGPT to evaluate decisions or draft reasoning autonomously. Training made judges more inclined to decide cases themselves and to use AI only for drafting their reasoning.

The researchers emphasize that their findings do not support the idea of replacing judges with AI. What the experiment shows, they say, is that AI can enhance public sector productivity, but only when paired with training that directs users toward the right tasks. Without this, most gains disappear.

The productivity gains in this study are likely a minimum, not a maximum. JudgeGPT operated on GPT-4, a pre-reasoning model that was best suited for the project at the time. Today's reasoning models write better, hallucinate less, and handle complex tasks more reliably.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.