The Crucial Role of Biological Data in AI Drug Discovery

GSK is expanding its AI-driven drug discovery efforts through a significant collaboration with Relation Therapeutics. This partnership leverages Relation’s expertise in generating high-volume, sophisticated human cell datasets. These datasets will train advanced AI models, including Relation’s MORGAN platform, to rapidly pinpoint novel drug targets. Building on prior successful collaborations, this agreement integrates cutting-edge data generation with AI model development for more precise therapeutic discovery.

Global pharmaceutical giant GSK is deepening its commitment to artificial intelligence-driven drug discovery through an expansive research collaboration with British biotech firm Relation Therapeutics. This strategic partnership, potentially valued at up to $110 million, signifies a significant expansion of their ongoing work, focusing on leveraging AI to unlock novel therapeutic avenues.

At the core of this collaboration lies Relation Therapeutics’ unique ability to generate high-volume, sophisticated datasets. These datasets will meticulously capture how human cells react to genetic alterations and diverse drug interventions. The rich biological information gleaned from these experiments will serve as the crucial training material for advanced AI models, including those built on Relation’s proprietary MORGAN platform. The objective is to empower these AI systems to pinpoint promising drug targets with unprecedented speed and accuracy.

This agreement intelligently intertwines the generation of granular biological data with cutting-edge AI model development. Relation Therapeutics’ innovative research methodology is characterized by its symbiotic integration of rigorous computational analysis with experimental approaches that yield fresh insights into cellular behavior. This holistic approach ensures that AI models are trained on data that is not only vast but also deeply biologically relevant.

The current initiative builds upon a foundation of prior successful collaborations between GSK and Relation, which previously focused on the challenging therapeutic areas of fibrotic diseases and osteoarthritis. Those earlier projects were instrumental in establishing functional disease datasets through observational studies, meticulously prepared for analysis using Relation’s sophisticated Lab-in-the-Loop platform. This prior work integrated a powerful combination of human genetics, multi-omics data derived from human tissues at the single-cell level, functional assays, and advanced machine learning techniques to identify and validate key disease targets.

Unlocking Biological Insights: Relation’s Data Generation Engine

Relation Therapeutics defines its distinctive Lab-in-the-Loop approach as a seamless fusion of laboratory experimentation and sophisticated computational analysis. Their expertise encompasses a wide array of cutting-edge techniques, including comprehensive tissue profiling, single-cell and spatial transcriptomics, advanced sequencing methodologies, and rigorous target validation. Concurrently, machine learning algorithms are strategically deployed for target identification, prioritization, validation, and the intelligent design of future experiments.

A key component of Relation’s methodology involves conducting carefully controlled perturbation experiments. These experiments meticulously measure how specific genetic modifications impact cellular characteristics directly associated with disease states. The resulting data is then analyzed in conjunction with existing genetic information and biological data derived from patient samples, offering a multi-dimensional view of disease mechanisms.

While public repositories of biological data remain invaluable resources for training foundational AI models, integrating information from disparate studies often presents considerable technical hurdles. A comprehensive review published in 2025 in Experimental & Molecular Medicine highlighted the immense value of repositories like CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus, which collectively provide researchers with access to vast quantities of single-cell data. CZ CELLxGENE alone, for instance, offers access to over 100 million standardized cells.

However, the diversity in sampling methods, sequencing protocols, experimental procedures, and data processing pipelines across different studies can introduce inconsistencies. Single-cell data is also susceptible to technical noise and various artifacts, necessitating meticulous dataset selection, rigorous filtering, careful composition balancing, and stringent quality control measures during the training of foundational models.

Furthermore, dataset overlap poses another significant challenge. The same or similar cell types can appear across multiple public resources, potentially leading to their disproportionate influence during model training. This overlap can also introduce data leakage risks when training and testing datasets share common elements.

The aforementioned review concluded that the meticulous assembly of high-quality, non-redundant datasets is as critical to building robust single-cell foundation models as the underlying model architecture itself.

The Nuances of Scale: Dataset Size and AI Model Performance

Research published in the prestigious journal Nature Methods in June of this year delved into the intricate relationship between the size and diversity of pretraining data and the performance of single-cell foundation models. Utilizing a corpus comprising 22.2 million cells, researchers trained an impressive 400 models and evaluated their efficacy across 6,400 distinct experiments.

The study’s findings revealed a surprising trend: current single-cell foundation models tended to reach performance plateaus after being trained on only a fraction of the available data. In contrast to the well-documented data-scaling laws observed in large language models, the AI systems examined in this study did not exhibit a consistent pattern where continually increasing training data directly translated into proportionally better results.

The researchers emphasized the need for a balanced approach, suggesting that model capacity, dataset size, and computational resources must be thoughtfully integrated rather than simply scaled up in unison. While the study did not definitively conclude that smaller or proprietary datasets are inherently superior, it strongly indicated that simply augmenting biological training data did not consistently yield further performance improvements.

A separate investigation published in Genome Biology in 2025 assessed two prominent single-cell foundation models, Geneformer and scGPT, across a range of zero-shot evaluation tasks. The study found that these models did not consistently outperform simpler analytical approaches. Moreover, the researchers identified challenges related to batch effects and cautioned against the assumption that larger, pretrained models automatically translate to superior biological representations.

The Strategic Pursuit of Specialized Datasets in Pharma

Relation Therapeutics has already demonstrated the power of its data-generation approach through Osteomics, which the company describes as a proprietary functional single-cell bone atlas. This ambitious project leverages patient-derived samples and integrates single-cell and spatial omics data with advanced imaging, genomics, proteomics, and detailed clinical phenotype information.

According to Relation Therapeutics, Osteomics is currently being employed to deepen the understanding of disease biology, identify novel therapeutic targets, discover predictive biomarkers, and stratify patient subgroups within the context of osteoporosis. The observational study benefits from the involvement of leading hospitals and research institutions across the UK and Australia.

Further underscoring the significance of this research area, a study published last month in Nature Genetics explored the cellular and genetic determinants of skeletal disease using a combination of single-cell analysis, genetic data, and functional validation. Notably, several researchers from Relation Therapeutics were co-authors on this impactful paper.

An analysis conducted in 2025 for Nature Biotechnology, focusing on AI-centric biopharma deals, identified specialized dataset providers as a prominent trend emerging from recent strategic partnerships. Other notable trends included larger upfront payments, the exploration of novel therapeutic modalities, and increased participation from established large biotechnology companies.

The Nature Biotechnology analysis posited that high-quality, disease-specific datasets are rapidly becoming a critical input for causal and generative machine-learning models. As an illustrative example, the analysis cited GSK’s separate agreement with Ochre Bio, valued at $37.5 million, which involved the licensing of human liver single-cell and perfused-organ data.

Another significant collaboration involved AstraZeneca and Pathos AI entering into a $200 million agreement with Tempus in 2025. Under this arrangement, Pathos was tasked with developing advanced oncology foundation models by utilizing de-identified clinical, genomic, and imaging data encompassing over 150,000 patients.

Despite these advancements, access to sufficient high-quality data continues to represent a significant constraint in the field of AI-driven drug discovery. A research highlight in Nature concerning federated learning in pharmaceutical research identified limited access to suitable training data as a major bottleneck for AI applications. Furthermore, companies often face restrictions on sharing proprietary information, adding another layer of complexity.

Consequently, AI-biopharma agreements exhibit considerable variability in how companies secure data and computational capabilities. Some collaborations are primarily centered on gaining access to existing AI platforms, while others encompass joint development initiatives, strategic data licensing agreements, or the co-creation of entirely new biological datasets.

The newly inked GSK–Relation Therapeutics agreement is particularly noteworthy for its comprehensive scope, encompassing both robust data generation and advanced model development. Relation Therapeutics is set to produce novel human cellular datasets as a direct outcome of this collaboration, which will then be instrumental in training sophisticated AI models designed to identify promising drug targets.

Original article, Author: Samuel Thompson. If you wish to reprint this article, please indicate the source:https://aicnbc.com/24336.html

Like (0)
Previous 2 days ago
Next 2 hours ago

Related News