Brief IA

Crosby Assesses Legal AI with Redline Bench

🛠️ AI Tools·Tom Levy·

Crosby Assesses Legal AI with Redline Bench

Crosby Assesses Legal AI with Redline Bench
Key Takeaways
1Crosby launches the Redline Bench, a tool to assess the effectiveness of AI in contract review.
2Crosby's benchmark allows for measuring the quality of contract modifications proposed by AI.
3ChatGPT 5.5 stands out with a score of 50.5% in the Redline Bench tests.
💡Why it mattersThis initiative could transform lawyers' trust in AI, impacting the costs and efficiency of legal services.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Crosby Innovates with an Assessment Tool for Legal AI

Crosby, a cutting-edge law firm, has recently unveiled an innovative tool designed to evaluate the performance of artificial intelligence models in the legal field. Named Redline Bench, this framework aims to help lawyers assess the reliability of AI technologies in contract negotiation.

An Industry Seeking Legitimacy

The legal tech sector is striving for recognition, but to achieve this, it must prove that its tools can think like lawyers. Crosby, positioning itself as both a startup and a law firm, provides legal services to companies like Cursor and Rogo. With the launch of Redline Bench, the firm aims to evaluate AI models on concrete legal tasks, starting with contract review.

AI Facing Legal Challenges

In recent years, engineers have observed that AI systems are becoming increasingly proficient in tasks such as coding. Legal tech companies now hope to develop AIs capable of reviewing contracts, identifying risks, and negotiating terms more effectively than human lawyers. However, the law presents unique challenges. Ryan Daniels, a former lawyer and founder of Crosby, emphasizes the difficulty of defining what is "good" or "bad" in legal work, unlike coding, where success is more easily measurable.

The Complexity of Legal Automation

This ambiguity complicates efforts to automate legal work. Companies like Anthropic have attempted to attract lawyers with specialized tools, drawing the attention of investors. However, the lack of a common framework to assess the quality of AI work remains a major obstacle.

Creating a New Benchmark

To address this need, Crosby has formed a team called Crosby Intelligence, composed of engineers and lawyers, to develop agents and an evaluation framework. Among them, Sharan Ramjee, a fraud detection expert, and Ross Weiser, a former lawyer at Sullivan & Cromwell, bring their expertise.

Collaboration with Micro1

Crosby has partnered with Micro1 to recruit lawyers capable of defining the criteria for good legal work. To build the framework, senior lawyers simulated transactions and rated crucial contract modifications, which were then transformed into weighted criteria.

A Rigorous Evaluation Process

During testing, AI models receive the same contracts used to establish the framework, and their modifications are compared by a panel of judges. The final score reflects how often the models make changes deemed important by the lawyers.

Transparency and Publication of Results

The Redline Bench will be accessible to all labs wishing to test their models according to Crosby's criteria. The company plans to regularly publish reports to compare the performances of leading models.

Promising Initial Results

In the first series of tests, ChatGPT 5.5 achieved a score of 50.5%, indicating that its modifications aligned with half of the lawyers' priorities. Gemini 3.5 Flash and Claude Opus 4.8 followed with scores of 45.1% and 44.4%, respectively.

A Promising Model Withdrawn

Crosby also tested Fable 5 from Anthropic, which received a promising score of 47.3% before being withdrawn from the market. Once access is restored, Crosby plans to reevaluate this model.

A Trend Towards Performance Measurement

Crosby is not alone in this endeavor. Harvey, another legal startup, along with Anthropic and OpenAI, are also developing their own frameworks to assess AI performance on concrete tasks. However, Ryan Daniels warns against the temptation for labs to adjust their systems to excel in their own tests.

Financial Stakes and Lawyer Trust

Beyond the scores, billions of dollars in investments hinge on AI's ability to reduce legal costs and lighten the workload of legal advisors. For lawyers to adopt these tools, they must trust in their reliability. Crosby is working to provide them with that assurance.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.