GSK and Relation: $110 Million for AI in Medicine

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A Strategic Partnership for Pharmaceutical Innovation
Pharmaceutical company GSK recently announced a strategic collaboration with Relation Therapeutics, a British biotechnology firm. This partnership, valued at $110 million, aims to enhance efforts in drug discovery assisted by artificial intelligence (AI). This initiative is part of a broader effort to integrate advanced technologies into the pharmaceutical development process.
As part of this agreement, Relation Therapeutics is committed to generating large-scale datasets. These data, crucial for research, measure human cell responses to genetic modifications and drug interventions. The goal is to train AI models capable of identifying new potential therapeutic targets. These models will be integrated into Relation's MORGAN platform, which is central to this collaboration.
The Importance of Biological Data in Research
This partnership highlights the central role of biological data in the development of AI models. Relation Therapeutics adopts an approach that combines computational analysis with laboratory experiments. This method allows for the generation of unprecedented insights into human cell behavior, paving the way for innovative medical discoveries.
Previous collaborations between GSK and Relation have focused on diseases such as fibrotic diseases and osteoarthritis. These projects involved observational studies aimed at creating functional datasets for in-depth analysis using Relation's Lab-in-the-Loop platform. These efforts have combined human genetics with single-cell multi-omic data derived from human tissues, functional assays, and machine learning to identify and validate new therapeutic targets.
Relation's Lab-in-the-Loop Method
Relation Therapeutics employs an innovative approach called Lab-in-the-Loop, which integrates laboratory experimentation with computational analysis. This method includes several advanced techniques such as tissue profiling, single-cell and spatial transcriptomics, sequencing, and target validation. Machine learning is at the core of this process, facilitating the identification, prioritization, and validation of targets as well as experimental design.
The company also conducts perturbation experiments to assess the impact of genetic modifications on disease-related cellular characteristics. These results are then analyzed in conjunction with biological data derived from patients and genetic information. Public repositories remain a valuable resource for training fundamental biological models, although combining data from different studies can present technical challenges.
Challenges Related to Single-Cell Data
A review from 2025 in Experimental & Molecular Medicine emphasized the importance of repositories such as CZ CELLxGENE, the Human Cell Atlas, and the NCBI Gene Expression Omnibus. These resources provide researchers with access to substantial volumes of single-cell data. For example, CZ CELLxGENE offers over 100 million standardized cells. However, differences in sampling methods, sequencing protocols, and experimental procedures across studies can introduce technical noise and other artifacts, necessitating rigorous selection of datasets.
Data overlap is another issue. The review noted that similar cells may appear in multiple public resources, risking bias in model training and creating data leaks when training and test datasets overlap. Thus, building robust foundational models requires not only a solid model architecture but also a high-quality, non-redundant dataset.
The Impact of Dataset Size
A study published in Nature Methods in June explored the effect of dataset size and diversity on single-cell foundational models. Using a corpus of 22.2 million cells, researchers trained 400 models and evaluated them across 6,400 experiments. The results showed that models often reach a performance plateau after being trained on a fraction of the available data, unlike large language models.
Researchers concluded that increasing dataset size does not necessarily guarantee better performance. It is crucial to find a balance between model capacity, data size, and computational resources. The study also revealed that smaller or proprietary datasets are not inherently superior, but adding additional biological training data does not systematically lead to performance gains.
The Importance of Specialized Datasets
Relation Therapeutics has already implemented its data generation strategy with Osteomics, a proprietary single-cell functional atlas of bone. This project utilizes patient-derived samples and integrates single-cell and spatial omics data with imaging, genomic, proteomic, and clinical phenotype information. Osteomics is used to explore disease biology, therapeutic targets, biomarkers, and patient subgroups in osteoporosis.
Recent research published in Nature Genetics also studied the cellular and genetic determinants of skeletal diseases using single-cell analysis, genetic data, and functional validation. Several researchers from Relation contributed to this study.
Emerging Trends in Biopharmaceutical Partnerships
A 2025 analysis from Nature Biotechnology highlighted emerging trends in biopharmaceutical agreements focused on AI. Specialized dataset providers have become key players, with larger upfront payments and increased participation from major biotechnology companies. The analysis emphasized that high-quality, disease-specific datasets are essential for causal and generative machine learning models.
GSK's distinct agreement with Ochre Bio, valued at $37.5 million, for licensing human liver single-cell data and perfused organs is one example. Similarly, AstraZeneca and Pathos AI signed a $200 million agreement with Tempus to develop foundational models in oncology using anonymized clinical, genomic, and imaging data covering over 150,000 patients.
Access to high-quality data remains a major challenge in AI-driven drug discovery. A Nature study identified limited access to appropriate training data as a significant barrier for AI applications, while noting that companies may also face restrictions on sharing proprietary information.
AI-biopharma agreements vary based on data acquisition methods and computational capabilities. Some focus on access to AI platforms, while others include joint development, data licensing, or the creation of new biological datasets. The agreement between GSK and Relation encompasses both data generation and model development, with Relation producing human cellular datasets to train AI models aimed at identifying new therapeutic targets.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.