AI Coalition releases predicted structures for 2,800-plus viral protein complexes

NVIDIA, Google DeepMind and EMBL-EBI open AlphaFold data and a GPU pipeline as U.N. talks turn to the next outbreak

Scientists got a lucky break with COVID-19. Years of earlier work on coronaviruses meant they already knew the key proteins well enough to design vaccines fast. The next pandemic might not be that generous.

NVIDIA has joined Google DeepMind, the European Molecular Biology Laboratory’s European Bioinformatics Institute and several other labs to put predicted 3D structures for the protein complexes of more than 2,800 viruses into the open AlphaFold Database. Any researcher can use them.

The structures were inferred with AlphaFold2, Google DeepMind’s model for how proteins fold. NVIDIA’s BioNeMo Inference Runtime handled the optimization so the team could run inference across thousands of viral proteomes and predict the complexes — the groups of proteins that actually work together inside each virus.

“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks.”

NVIDIA is also releasing the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow that produced the set. Researchers can take their own protein sequences and generate predicted structures without starting from scratch.

Preparation can’t wait. An analysis by the Center for Global Development puts the chance of another pandemic as severe as COVID-19 at roughly 50 percent by 2050.

“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is stockpile some of that knowledge ahead of time.”

About 30 percent of the protein interactions now being added have never been seen. Their shapes do not appear in the Protein Data Bank, the main store of experimentally solved structures. That is new ground for anyone hunting targets.

“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward.”

Most proteins do not work alone. They assemble into complexes that carry out the real work. Those assemblies are often what a vaccine or a drug has to hit. The spike protein on the COVID-19 virus is the obvious example. Its structure made vaccine design possible. For thousands of other viruses, that map simply does not exist. This release starts to fill the hole.

Old methods — grow crystals, fire X-rays — can take years and cost thousands of dollars per structure. AlphaFold2, tuned with NVIDIA BioNeMo to run on NVIDIA GPUs, produces a prediction in minutes and can be batched. High-confidence hits can still be checked in the lab.

The team worked through viral families known to infect people, from ordinary cold viruses to newer threats such as Mpox.

Partners include the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.

The data dropped the same week a United Nations General Assembly meeting convened by the World Economic Forum is taking up pandemic prevention, preparedness and response in New York. The AlphaFold Database now holds more than 260 million protein and protein-complex predictions, covering nearly every cataloged protein known to science.

“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI. “The dataset also covers lesser-studied viruses and lowers the barriers for scientists in low-resource settings who are confronting outbreaks firsthand.”

Each prediction carries a confidence score. The models show how a viral complex might look and how its proteins could interact.

“When I did my Ph.D., there were no structures for any of the proteins we were investigating. It was like working in the dark — we had to guess what was going on,” Grove said. “This dataset is a powerful tool for all the researchers doing their Ph.D.s now, giving them high-quality structural data that’s going to accelerate fundamental science.”

The viral complexes are on the AlphaFold Database Pandemic Preparedness Portal. The prediction pipeline is on GitHub.


Subscribe
Notify of
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x