Yovao News · The World, In Focus. From Local to Global, Never Miss a Beat

Proprietary Pharma Data Drives Leap in AI Protein Modeling

Proprietary Pharma Data Drives Leap in AI Protein Modeling

An AI system trained on more than 20,000 proprietary protein structures from pharmaceutical companies has surpassed the performance of AlphaFold-like models that rely exclusively on publicly available data. The findings, detailed in a recent blog post by a consortium of drug firms, suggest that internal industry data is a critical resource for advancing protein-folding technology.

Although the study has not undergone peer review and the resulting model remains unavailable to the public, the results highlight a growing consensus that public databases alone are insufficient for the next generation of drug discovery tools. Researchers argue that incorporating data on how proteins interact with potential drugs is essential for improving model accuracy.

The group utilized OpenFold3, an open-source replication of Google DeepMind’s AlphaFold 3, to develop the new model. It was fine-tuned using 20,167 structures captured from the internal vaults of companies including AbbVie and Astex Pharmaceuticals. These structures, generated through techniques such as X-ray crystallography and cryo-electron microscopy, represent protein-ligand interactions that are rarely deposited in public repositories because they pertain to proprietary development efforts.

“You add all this data, and you get a pretty big bump in performance,” said Mohammed AlQuraishi, a computational biologist at Columbia University in New York who participated in the project. The new model outperformed both comparable systems trained solely on public data and those trained on the siloed datasets of individual firms.

The initiative stems from a limitation in existing protein-folding tools. The Protein Data Bank (PDB), which hosts over 200,000 experimentally determined structures, served as the foundation for AlphaFold 2 and its subsequent Nobel Prize recognition. However, Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals, estimates that the PDB contains only about 10,000 examples of structures interacting with drug-like molecules.

This scarcity becomes a significant bottleneck for drug discovery. Research indicates that the accuracy of “co-folding” models, which predict how proteins interact with other molecules, drops sharply when challenged with molecular pairs that differ significantly from their training data.

To address this, the AI Structural Biology (AISB) Network was formed last year by AbbVie, Astex, and other pharmaceutical partners. The collaboration aims to bridge the gap between academic open science and industry secrecy. AlQuraishi noted that these findings strengthen the case for creating similar publicly accessible datasets, pointing to the UK-funded OpenBind project, which recently released hundreds of new structures with thousands more expected.

5 responses to “Proprietary Pharma Data Drives Leap in AI Protein Modeling”

  1. Finally some proof that corporate secrets can drive real science forward when shared responsibly among collaborators like this.

  2. So we are just creating more data silos now? I worry this widens the gap between well-funded giants and smaller labs.

  3. It makes perfect sense. Protein-ligand interactions are too commercially sensitive for most pharma companies to release into the PDB.

  4. Wait, the model isn’t even public yet? How can the scientific community verify these impressive performance claims without peer review?

Leave a Reply

Your email address will not be published. Required fields are marked *