Sony Pixel Power calrec Sony

Speaking the Language of the Genome: Gordon Bell Finalist Applies Large Language Models to Predict New COVID Variants

14/11/2022

A finalist for the Gordon Bell special prize for high performance computing-based COVID-19 research has taught large language models (LLMs) a new lingo - gene sequences - that can unlock insights in genomics, epidemiology and protein engineering.

Published in October, the groundbreaking work is a collaboration by more than two dozen academic and commercial researchers from Argonne National Laboratory, NVIDIA, the University of Chicago and others.

The research team trained an LLM to track genetic mutations and predict variants of concern in SARS-CoV-2, the virus behind COVID-19. While most LLMs applied to biology to date have been trained on datasets of small molecules or proteins, this project is one of the first models trained on raw nucleotide sequences - the smallest units of DNA and RNA.

We hypothesized that moving from protein-level to gene-level data might help us build better models to understand COVID variants, said Arvind Ramanathan, computational biologist at Argonne, who led the project. By training our model to track the entire genome and all the changes that appear in its evolution, we can make better predictions about not just COVID, but any disease with enough genomic data.

The Gordon Bell awards, regarded as the Nobel Prize of high performance computing, will be presented at this week's SC22 conference by the Association for Computing Machinery, which represents around 100,000 computing experts worldwide. Since 2020, the group has awarded a special prize for outstanding research that advances the understanding of COVID with HPC.

Training LLMs on a Four-Letter Language LLMs have long been trained on human languages, which usually comprise a couple dozen letters that can be arranged into tens of thousands of words, and joined together into longer sentences and paragraphs. The language of biology, on the other hand, has only four letters representing nucleotides - A, T, G and C in DNA, or A, U, G and C in RNA - arranged into different sequences as genes.

While fewer letters may seem like a simpler challenge for AI, language models for biology are actually far more complicated. That's because the genome - made up of over 3 billion nucleotides in humans, and about 30,000 nucleotides in coronaviruses - is difficult to break down into distinct, meaningful units.

When it comes to understanding the code of life, a major challenge is that the sequencing information in the genome is quite vast, Ramanathan said. The meaning of a nucleotide sequence can be affected by another sequence that's much further away than the next sentence or paragraph would be in human text. It could reach over the equivalent of chapters in a book.

NVIDIA collaborators on the project designed a hierarchical diffusion method that enabled the LLM to treat long strings of around 1,500 nucleotides as if they were sentences.

Standard language models have trouble generating coherent long sequences and learning the underlying distribution of different variants, said paper co-author Anima Anandkumar, senior director of AI research at NVIDIA and Bren professor in the computing + mathematical sciences department at Caltech. We developed a diffusion model that operates at a higher level of detail that allows us to generate realistic variants and capture better statistics.

Predicting COVID Variants of Concern Using open-source data from the Bacterial and Viral Bioinformatics Resource Center, the team first pretrained its LLM on more than 110 million gene sequences from prokaryotes, which are single-celled organisms like bacteria. It then fine-tuned the model using 1.5 million high-quality genome sequences for the COVID virus.

By pretraining on a broader dataset, the researchers also ensured their model could generalize to other prediction tasks in future projects - making it one of the first whole-genome-scale models with this capability.

Once fine-tuned on COVID data, the LLM was able to distinguish between genome sequences of the virus' variants. It was also able to generate its own nucleotide sequences, predicting potential mutations of the COVID genome that could help scientists anticipate future variants of concern.

Trained on a year's worth of SARS-CoV-2 genome data, the model can infer the distinction between various viral strains. Each dot on the left corresponds to a sequenced SARS-CoV-2 viral strain, color-coded by variant. The figure on the right zooms into one particular strain of the virus, which captures evolutionary couplings across the viral proteins specific to this strain. Image courtesy of Argonne National Laboratory's Bharat Kale, Max Zvyagin and Michael E. Papka. Most researchers have been tracking mutations in the spike protein of the COVID virus, specifically the domain that binds with human cells, Ramanathan said. But there are other proteins in the viral genome that go through frequent mutations and are important to understand.

The model could also integrate with popular protein-structure-prediction models like AlphaFold and OpenFold, the paper stated, helping researchers simulate viral structure and study how genetic mutations impact a virus' ability to infect its host. OpenFold is one of the pretrained language models included in the NVIDIA BioNeMo LLM service for developers applying LLMs to digital biology and chemistry applications.

Supercharging AI Training With GPU-Accelerated Supercomputers The team developed its AI models on supercomputers powered by NVIDIA A100 Tensor Core GPUs - including Argonne's Polaris, the U.S. Department of Energy's Perlmutter, and NVIDIA's in-house Selene system. By scaling up to these powerful systems, they achieved performance of more than 1,500 exaflops in training runs, creating the largest biological language models to date.

We're working with models today that have up to 25 billion
LINK: https://blogs.nvidia.com/blog/2022/11/14/genomic-large-language-model-...
See more stories from nvidia

Most recent headlines

01/12/2025

L3Harris and PentenAmio Formalise Agreement to Advance Key Management and Secure Communications Technology

L3Harris and PentenAmio formalise their teaming agreement at MilCIS 2025, streng...

01/12/2025

Artemis II: A Mission of Veterans, Firsts and Lunar Dreams

Artemis II is NASA's first crewed flight test of the Space Launch System rocket and Orion spacecraft. The crew, from left: Commander Reid Wiseman, Pilot Vic...

01/12/2025

Wooden Camera Releases Accessory Collection for Canon EOS C50

IRVINE, Calif. Wooden Camera has introduced its new Accessory Collection for the Canon EOS C50. The new lineup includes a low-profile, gimbal-ready cage, expand...

01/12/2025

FCC to Vote on LPTV Rules at December Public Meeting

WASHINGTON The Federal Communications Commission has released a tentative agenda for its Dec. 18 Open Commission Meeting that will include a vote on a report an...

01/12/2025

2026 Local TV Ad Forecasts Offer Growth and Uncertainties

In most years, a graph of annual local TV ad spending is about as predictable as an electrocardiogram of a reasonably healthy patient in a doctor's office. ...

01/12/2025

Increasingly Software-Centric Switchers Occupy Hybrid Space

Many industries have seen big-ticket hardware turn into software. Switchers, though, demand a combination of real-time performance and sheer bandwidth that has ...

01/12/2025

China to Host ITU World Radiocommunication Conference 2027

GENEVA Shanghai will host the next quadrennial Radiocommunication Assembly (RA-27) and World Radiocommunication Conference (WRC-27), Oct. 11-Nov. 12, 2027. This...

01/12/2025

Broadcasters Foundation Seeks Donations for Giving Tuesday

NEW YORK Just in time for Giving Tuesday tomorrow (Dec. 2), the Broadcasters Foundation of America is seeking out donations to help television and radio industr...

01/12/2025

Net Insight CEO Crister Fritzson Sets 2026 Retirement

STOCKHOLM, Sweden Net Insight CEO Crister Fritzson has informed the company's board that he will retire from the video transport and media cloud technology ...

01/12/2025

Kyivstar and Ukrainian Ministry of Digital Transformation Select Google Gemma as the Foundation for Ukraine's National LLM

01 Dec 2025 Kyivstar and Ukrainian Ministry of Digital Transformation Select Go...

01/12/2025

Sky unveils official trailer for highly anticipated prequel Gomorrah The Origins coming to Sky and NOW early 2026

The prequel to the global hit Sky Original mob crime saga, Gomorrah' is a s...

01/12/2025

Step inside Nick Caves Veiled World a one-off special featuring Florence Welch, Flea and more, premiering 6 December on Sky and NOW

Featuring Florence Welch, Red Hot Chilli Peppers' Flea, designer Bella Freud...

01/12/2025

A DowntoEarth, AllTooRelatable Hero: Cashero' Teaser Trailer Unveiled, Premieres December 26

Back to All News A Down to Earth, All Too Relatable Hero: Cashero' Teaser ...

01/12/2025

The Hidden Impact of Bad Conversion

In the rush to deliver content to every screen, many broadcasters overlook one of the most crucial steps in the workflow: format and frame rate conversion. Get ...

01/12/2025

Fox Corporation Chief Financial Officer Steve Tomsic to Participate in Upcoming UBS Global Media and Communications Conference 2025

Fox Corporation Chief Financial Officer Steve Tomsic to Participate in Upcoming ...

01/12/2025

Arvato Systems' Virtual Private Cloud (VPC) Receives BSI C5 Certification (Type 2) Following Audit by HLB Stckmann

Arvato Systems' Virtual Private Cloud (VPC) Receives BSI C5 Certification (T...

01/12/2025

Consumer Magazines Returns System - 2025 Update

Summary This short video gives you an summary of the changes in under two minutes. --...

01/12/2025

From Ballet to Books: RT is Supporting 12 Arts and Cultural Events all over Ireland this December

As the festive season approaches, RT Supporting the Arts is proud to showcase a...

01/12/2025

At NeurIPS, NVIDIA Advances Open Model Development for Digital and Physical AI

Researchers worldwide rely on open-source technologies as the foundation of their work. To equip the community with the latest advancements in digital and physi...

01/12/2025

Lights, Camera, Christmas: RT rings in the Season S an Nollaig

Festive specials of Christmas in Kilmainham presented by Marty Whelan, High Road Low Road, Callan Kicks the Year and Keys to My Life Ring in the New Year with ...

01/12/2025

Architect and presenter Hugh Wallace dies aged 68

Architect and television presenter Hugh Wallace, best known to RT audiences as a long-serving judge on Home of the Year, has died at the age of 68. In a state...

28/11/2025

Brides Asks for Compassion for Our Youths

Nadia Fall attends the 2025 Sundance Film Festival premiere of Brides at the Egyptian Theatre on January 24, 2025, in Park City, Utah. (Photo by Donyale West/...

28/11/2025

4 Reasons Why Keeping Your Spotify App Updated Matters and What You Might Be Missing

It's easy to ignore those little red update available badges. But when it ...

28/11/2025

FCC to Vote on LPTV Rules at Dec. Public Meeting

WASHINGTON Federal Communications Commission has released a tentative agenda for the December Open Commission Meeting scheduled for Thursday, December 18, 2025 ...

28/11/2025

Professional Fighters League Packs a Domestic, International MMA Punch (TV Sportsplay)

The Professional Fighters League is looking to super-serve fans of mixed martial...

28/11/2025

Fubo Launches Multiview Beta on Roku

Fubo has released in beta on select Roku devices a new feature that lets users display up to four simultaneous streams at once....

28/11/2025

WNBA Playoffs Continue: What's On This Weekend in TV Sports (Sept. 28-29)

The WNBA playoffs and Week 4 of the NFL regular season highlight the list of live sports events airing on television this weekend....

28/11/2025

Freeze Frame: B+C Hall of Fame 2024

The 32nd class of honorees to the B+C Hall of Fame took to the stage at New York's Ziegfeld Ballroom on September 26 for a gala induction event. Click below...

28/11/2025

Next Text: As DirecTV and Dish Try to Seize the Remains of the Day, Does It Even Matter?

We hold in our hands the very last Next Text for Next TV, the weekly back-and-fo...

28/11/2025

DirecTV Acquires Dish, Unifying Struggling Satellite Business

DirecTV said it made a deal with EchoStar to buy EchoStar's video businesses, including satellite-TV provider Dish TV and virtual MVPD Sling TV, for $1 plus...

28/11/2025

B+C Hall of Fame Announces Its Class of 2025

The Broadcasting+Cable Hall of Fame, the premier industry event paying tribute to the influencers, innovators and shining lights of broadcast, cable and streami...

28/11/2025

Sky Sports x Slawn drop limited-edition football jersey that unlocks a month of free content from the home of sport

Friday 28 November 2025 Sky Sports x Slawn drop limited-edition football jersey...

28/11/2025

Rohde & Schwarz shows resilience in a challenging environment, revenue exceeds three billion euros for the first time

Rohde & Schwarz shows resilience in a challenging environment, revenue exceeds t...

28/11/2025

Changing children's lives for good: Donations for the RT Toy Show Appeal 2025 open tonight

Unwrapped: The Toy Show Appeal - airing this Sunday on RT One and RT Player- s...

27/11/2025

Vizrt Launches Viz One 8.1 With AI-Powered Features

LONDON Vizrt has added several AI-driven advanced features offering improved speed, intelligence and accuracy in the newest version of its media asset managemen...

27/11/2025

Prime Video Debuts AI-Powered Video Recaps

Prime Video has launched AI-powered video season recaps in a beta version for select English-language Prime Original series in the U.S., a move Amazon is callin...

27/11/2025

Netflix's 'Raat Akeli Hai: The Bansal Murders' Marks a Grand World Premiere at IFFI Ahead of Its Global Release on 19th December

Back to All News Netflix's Raat Akeli Hai: The Bansal Murders Marks a Grand...

27/11/2025

Sky unveils first look image from high-stakes action thriller Prisoner, coming 2026

Tahar Rahim and Izuka Hoyle star in the gripping six-part Sky Original from Acad...

27/11/2025

Sky Arts Reveals the Nations Greatest Basslines and Queen Reign Supreme

Thursday 27 November 2025 Sky Arts Reveals the Nation's Greatest Basslines - and Queen Reign Supreme The UK's most iconic basslines have been revealed...

27/11/2025

Stranger Things 5': Prepare for One Last Adventure With Our Final Season Coverage Guide

Back to All News Stranger Things 5': Prepare for One Last Adventure With O...

27/11/2025

Elastic Compute for a Sustainable Media Industry

The media industry has a paradox at its core. It's an industry built on light, color and imagination, yet behind the scenes, it's powered by one of the ...

27/11/2025

Arqiva Achieves Five-Star GRESB Rating

Rating reflects rating progress across areas including policies, diversity & inclusion, health & safety and Net Zero leadership Winchester, UK, 27 November 202...

27/11/2025

Retail Media Audits Explained: What Networks Need to Know

What are the industry standards for Retail Media? Kathryn explains that certification is based on the IAB Europe Retail Media Measurement Standards and the IAB ...

27/11/2025

Katie Taylor, Rachael Blackmore and Arthur Gourounlian among the guests on this week's Late Late Show

World champion boxer and Irish sporting icon Katie Taylor will be in studio this...

27/11/2025

Tonight on RT Prime Time, serious child protection concerns emerge over online gaming platform, Roblox

Roblox, one of the world's most popular online gaming platforms for primary ...

27/11/2025

The Ultimate Black Friday Deal Is Here

Black Friday is leveling up. Get ready to score one of the biggest deals of the season - 50% off the first three months of a new GeForce NOW Ultimate membership...

26/11/2025

SVG Sit-Down: Prime Video EP Mike Muriano Previews Massive Black Friday Slate Featuring NFL, NBA, and Golf

SVG Sit-Down: Prime Video EP Mike Muriano Previews Massive Black Friday Slate Fe...