Brief IA

AI Revolutionizes Drug Discovery

🔬 Research·Tom Levy·

AI Revolutionizes Drug Discovery

AI Revolutionizes Drug Discovery
Key Takeaways
1Drug discovery is expensive and risky, with costs reaching up to $2.5 billion.
2AI promises to reduce timelines and improve success rates by quickly identifying new compounds.
3Quality data is crucial for the effectiveness of AI models, but biases and manipulation pose challenges.
💡Why it mattersAI could transform pharmaceutical research, but it requires reliable data to realize its potential.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI at the Heart of the Pharmaceutical Revolution

Drug discovery is a complex and costly process, marked by high risks. Since the 1950s, the cost of developing new treatments has nearly doubled every nine years, a phenomenon known as Eroom's Law. Currently, bringing a new drug to market takes between 10 and 15 years and can cost between $1 billion and $2.5 billion. Despite these colossal investments, the failure rate often exceeds 90%.

In the face of these challenges, the pharmaceutical industry is betting on artificial intelligence (AI) to improve success rates and reduce timelines. AI enables companies to better identify, test, and optimize new chemical compounds, thereby reducing the risk of costly failures at advanced stages of development.

Paul Belcher, Director of Protein Research Strategy at Cytiva, emphasizes that the clinical phase remains the most expensive in drug discovery. AI is seen as a crucial tool for mitigating risks and increasing the chances of success at this critical stage. It promises not only to save time but also to allow better candidates to reach clinical trials.

However, the use of AI in drug discovery highlights the need for robust and authentic data, as well as their integration into laboratory systems.

AI Transforms Laboratories

One of the most promising applications of AI in drug discovery is target identification. This process involves testing libraries of molecular entities against a disease-related target, such as a protein, to identify molecules capable of binding to it. Success at this stage provides researchers with a starting point for further testing and refinement, with the goal of developing a viable drug.

Belcher has observed a shift from empirical screening to predictive design. Rather than physically testing libraries, pharmaceutical companies are now using AI to design drug candidates from scratch and predict their interactions with disease targets before committing to research and development (R&D).

This means that companies are no longer limited by the number of physical tests they can conduct to identify starting points. “AI removes that constraint,” explains Belcher. “It also allows for the elimination of low-quality candidates before they are physically tested, thus saving time and resources.”

However, AI cannot yet reliably predict the kinetics or developability of new compounds, Belcher points out. Each candidate generated by AI still needs to be validated in the lab.

Traditional screening workflows were designed to identify targets at scale, not to profile a large number of complex candidates in detail. This places increased pressure on laboratory teams, who must test, characterize, and purify an ever-growing volume of AI-generated compounds.

“Current techniques used in target identification can test hundreds of thousands, even millions of compounds, using binary or threshold-based techniques that produce low-fidelity data — yes or no responses,” explains Belcher. “AI can increase the number of targets you obtain and potentially give you better targets as well. This increases the demand for high-throughput and information-rich technologies to validate and characterize these targets.”

The Crucial Importance of Quality Data

While AI has accelerated the demand for data-rich laboratory systems, it has also highlighted a fundamental need for better, more comprehensive data.

Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. As models have access to the same data, they all arrive at similar conclusions, with diminishing returns over time. Furthermore, the datasets were not built with AI in mind, meaning they lack the structure, labeling, and diversity necessary to keep models accurate and free from bias.

Publication bias exacerbates the problem. “Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.”

The data that Belcher believes could significantly improve models — failed experiments, compounds that do not bind — remain frustratingly difficult to obtain. “We often joke that there should be a negative data journal,” he says. “They are often buried in lab notebooks and never used to inform or guide future research.”

This absence of negative data creates a fundamental problem: without access to a wide range of data, models cannot be properly trained to avoid bias. “In all machine learning applications, the performance of the model heavily depends on the quality and extent of the training data,” notes Belcher.

Data fabrication has also become much easier with AI, raising concerns about data integrity. Take Western blots, for example. These are elements of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that about 4% of biomedical articles contained duplicated or manipulated images. This dates back to 2016, before generative AI made fabrication trivial.

“Manipulated or falsified data have always been a problem in science, but in the world of AI, especially when used to train models, this could have potentially disastrous consequences,” says Belcher. “There must be more tools to verify that the data is not manipulated.”

Some providers are beginning to tackle this challenge. Belcher mentions solutions like Cytiva's Image Integrity Checker, which uses secure hashing algorithms — the same technology used in blockchain — to detect whether scientific images have been altered. “We are starting to see a lot of interest from publishers who want to adopt this as a standard, as it is a quick way to ensure that what is published in the literature is authentic,” he adds.

Towards Autonomous Laboratories

Belcher describes the future state of drug discovery as fully autonomous laboratories that operate with minimal human intervention. Central to this vision is the consistency of data and infrastructure.

These AI-driven dark labs, or loop laboratories, operate 24/7. They go through cycles of prediction, testing, and optimization, then feed the results back into AI models to guide the next round of experiments. This could improve the success rates of drug candidates entering clinical trials, says Belcher. Better starting points, combined with more optimization cycles, should lead to better candidates with fewer risks reaching the clinic.

But automating a laboratory heavily depends on integration. This means interoperable systems, highly structured and comprehensive datasets, and smooth information flow. Most laboratories are not yet at this stage. “Today, many of the instruments in laboratories are standalone,” notes Belcher. “You can have the best technology in the world, but if it’s a closed ecosystem — if the user cannot extract the data — it’s worthless.”

An integrated infrastructure can enable laboratories to generate FAIR (Findable, Accessible, Interoperable, and Reusable) data at scale. This would not only inform individual laboratory reports but could also train the next generations of AI models, effectively closing the loop between the AI-driven dry lab and the physical wet lab.

“Our goal is to help scientists and researchers accelerate their breakthroughs and make this future state of autonomous laboratories a reality,” says Belcher. “We want to help them generate reliable data, streamline workflows in discovery, and hopefully what they work on becomes the groundbreaking therapies of tomorrow, faster and with more confidence.”

Financial Challenges and the Future of AI in Pharmacy

AI-driven drug discovery is still in its infancy. Notably, no drug discovered primarily through AI-driven design has yet received full approval from the FDA — although Belcher expects this to change in the next two to three years.

What impact could AI ultimately have on drug discovery? “The holy grail would be the complete in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work,” says Belcher. But there are many obstacles to this beyond model maturity, including regulatory hurdles and cost challenges.

A Stanford study found that the cost of training cutting-edge AI models has more than doubled each year since 2016, adding additional financial pressure to an industry already defined by exceptionally high R&D expenditures.

Belcher acknowledges the tension but remains optimistic about the future. “I think we will reach a point where there will be a balance between AI and lab work, from a cost and risk perspective,” he says. “As long as the cost of computation never exceeds the cost of clinical development, I believe AI will be an advantage.”

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.