Crosby Assesses Legal AI with Redline Bench

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Crosby Innovates with an Assessment Tool for Legal AI
Crosby, a cutting-edge law firm, has recently unveiled an innovative tool designed to evaluate the performance of artificial intelligence models in the legal field. Named Redline Bench, this framework aims to help lawyers assess the reliability of AI technologies in contract negotiation.
An Industry Seeking Legitimacy
The legal tech sector is striving for recognition, but to achieve this, it must prove that its tools can think like lawyers. Crosby, positioning itself as both a startup and a law firm, provides legal services to companies like Cursor and Rogo. With the launch of Redline Bench, the firm aims to evaluate AI models on concrete legal tasks, starting with contract review.
AI Facing Legal Challenges
In recent years, engineers have observed that AI systems are becoming increasingly proficient in tasks such as coding. Legal tech companies now hope to develop AIs capable of reviewing contracts, identifying risks, and negotiating terms more effectively than human lawyers. However, the law presents unique challenges. Ryan Daniels, a former lawyer and founder of Crosby, emphasizes the difficulty of defining what is "good" or "bad" in legal work, unlike coding, where success is more easily measurable.
The Complexity of Legal Automation
This ambiguity complicates efforts to automate legal work. Companies like Anthropic have attempted to attract lawyers with specialized tools, drawing the attention of investors. However, the lack of a common framework to assess the quality of AI work remains a major obstacle.
Creating a New Benchmark
To address this need, Crosby has formed a team called Crosby Intelligence, composed of engineers and lawyers, to develop agents and an evaluation framework. Among them, Sharan Ramjee, a fraud detection expert, and Ross Weiser, a former lawyer at Sullivan & Cromwell, bring their expertise.
Collaboration with Micro1
Crosby has partnered with Micro1 to recruit lawyers capable of defining the criteria for good legal work. To build the framework, senior lawyers simulated transactions and rated crucial contract modifications, which were then transformed into weighted criteria.
A Rigorous Evaluation Process
During testing, AI models receive the same contracts used to establish the framework, and their modifications are compared by a panel of judges. The final score reflects how often the models make changes deemed important by the lawyers.
Transparency and Publication of Results
The Redline Bench will be accessible to all labs wishing to test their models according to Crosby's criteria. The company plans to regularly publish reports to compare the performances of leading models.
Promising Initial Results
In the first series of tests, ChatGPT 5.5 achieved a score of 50.5%, indicating that its modifications aligned with half of the lawyers' priorities. Gemini 3.5 Flash and Claude Opus 4.8 followed with scores of 45.1% and 44.4%, respectively.
A Promising Model Withdrawn
Crosby also tested Fable 5 from Anthropic, which received a promising score of 47.3% before being withdrawn from the market. Once access is restored, Crosby plans to reevaluate this model.
A Trend Towards Performance Measurement
Crosby is not alone in this endeavor. Harvey, another legal startup, along with Anthropic and OpenAI, are also developing their own frameworks to assess AI performance on concrete tasks. However, Ryan Daniels warns against the temptation for labs to adjust their systems to excel in their own tests.
Financial Stakes and Lawyer Trust
Beyond the scores, billions of dollars in investments hinge on AI's ability to reduce legal costs and lighten the workload of legal advisors. For lawyers to adopt these tools, they must trust in their reliability. Crosby is working to provide them with that assurance.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.