
A finalist for the Gordon Bell special prize for high performance computing-based COVID-19 research has taught large language models (LLMs) a new lingo - gene sequences - that can unlock insights in genomics, epidemiology and protein engineering.
Published in October, the groundbreaking work is a collaboration by more than two dozen academic and commercial researchers from Argonne National Laboratory, NVIDIA, the University of Chicago and others.
The research team trained an LLM to track genetic mutations and predict variants of concern in SARS-CoV-2, the virus behind COVID-19. While most LLMs applied to biology to date have been trained on datasets of small molecules or proteins, this project is one of the first models trained on raw nucleotide sequences - the smallest units of DNA and RNA.
We hypothesized that moving from protein-level to gene-level data might help us build better models to understand COVID variants, said Arvind Ramanathan, computational biologist at Argonne, who led the project. By training our model to track the entire genome and all the changes that appear in its evolution, we can make better predictions about not just COVID, but any disease with enough genomic data.
The Gordon Bell awards, regarded as the Nobel Prize of high performance computing, will be presented at this week's SC22 conference by the Association for Computing Machinery, which represents around 100,000 computing experts worldwide. Since 2020, the group has awarded a special prize for outstanding research that advances the understanding of COVID with HPC.
Training LLMs on a Four-Letter Language LLMs have long been trained on human languages, which usually comprise a couple dozen letters that can be arranged into tens of thousands of words, and joined together into longer sentences and paragraphs. The language of biology, on the other hand, has only four letters representing nucleotides - A, T, G and C in DNA, or A, U, G and C in RNA - arranged into different sequences as genes.
While fewer letters may seem like a simpler challenge for AI, language models for biology are actually far more complicated. That's because the genome - made up of over 3 billion nucleotides in humans, and about 30,000 nucleotides in coronaviruses - is difficult to break down into distinct, meaningful units.
When it comes to understanding the code of life, a major challenge is that the sequencing information in the genome is quite vast, Ramanathan said. The meaning of a nucleotide sequence can be affected by another sequence that's much further away than the next sentence or paragraph would be in human text. It could reach over the equivalent of chapters in a book.
NVIDIA collaborators on the project designed a hierarchical diffusion method that enabled the LLM to treat long strings of around 1,500 nucleotides as if they were sentences.
Standard language models have trouble generating coherent long sequences and learning the underlying distribution of different variants, said paper co-author Anima Anandkumar, senior director of AI research at NVIDIA and Bren professor in the computing + mathematical sciences department at Caltech. We developed a diffusion model that operates at a higher level of detail that allows us to generate realistic variants and capture better statistics.
Predicting COVID Variants of Concern Using open-source data from the Bacterial and Viral Bioinformatics Resource Center, the team first pretrained its LLM on more than 110 million gene sequences from prokaryotes, which are single-celled organisms like bacteria. It then fine-tuned the model using 1.5 million high-quality genome sequences for the COVID virus.
By pretraining on a broader dataset, the researchers also ensured their model could generalize to other prediction tasks in future projects - making it one of the first whole-genome-scale models with this capability.
Once fine-tuned on COVID data, the LLM was able to distinguish between genome sequences of the virus' variants. It was also able to generate its own nucleotide sequences, predicting potential mutations of the COVID genome that could help scientists anticipate future variants of concern.
Trained on a year's worth of SARS-CoV-2 genome data, the model can infer the distinction between various viral strains. Each dot on the left corresponds to a sequenced SARS-CoV-2 viral strain, color-coded by variant. The figure on the right zooms into one particular strain of the virus, which captures evolutionary couplings across the viral proteins specific to this strain. Image courtesy of Argonne National Laboratory's Bharat Kale, Max Zvyagin and Michael E. Papka. Most researchers have been tracking mutations in the spike protein of the COVID virus, specifically the domain that binds with human cells, Ramanathan said. But there are other proteins in the viral genome that go through frequent mutations and are important to understand.
The model could also integrate with popular protein-structure-prediction models like AlphaFold and OpenFold, the paper stated, helping researchers simulate viral structure and study how genetic mutations impact a virus' ability to infect its host. OpenFold is one of the pretrained language models included in the NVIDIA BioNeMo LLM service for developers applying LLMs to digital biology and chemistry applications.
Supercharging AI Training With GPU-Accelerated Supercomputers The team developed its AI models on supercomputers powered by NVIDIA A100 Tensor Core GPUs - including Argonne's Polaris, the U.S. Department of Energy's Perlmutter, and NVIDIA's in-house Selene system. By scaling up to these powerful systems, they achieved performance of more than 1,500 exaflops in training runs, creating the largest biological language models to date.
We're working with models today that have up to 25 billion
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
06/09/2026
June 9 2026, 23:00 (PDT) Dolby and MagentaTV Bring Fans Closer to the FIFA Worl...
04/08/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
04/07/2026
April 7 2026, 19:00 (PDT) Detective Conan: Fallen Angel of the Highway Opens in...
15/06/2026
One of the more exciting internal video production divisions within a college at...
15/06/2026
The deal valued at $22 Billion is expected to close in the first half of 2027...
15/06/2026
Golf Channel and the Arnold Palmer Cup have announced a partnership to livestream the 2026 Arnold Palmer Cup on Golf Channel Mobile and GolfChannel.com. The tou...
15/06/2026
TikTok and Panini have announced a partnership to bring a digital collectible ca...
15/06/2026
Cosm and Monster Energy have announced the debut of the first full-dome immersiv...
15/06/2026
Real American Freestyle (RAF) and Fox Nation have announced an exclusive streaming agreement for three RAF international events, beginning with RAF Georgia on J...
15/06/2026
FanConnect has announced a partnership with Extreme Networks integrating FanConn...
15/06/2026
Ten Emerging Filmmakers Ages 18 to 25 Will Start Fellowship Year at Ignite Lab from June 14-19
LOS ANGELES, CA, June 15, 2026 - The nonprofit Sundance Institut...
15/06/2026
Innovative three-band soft synth introduced
UVI's latest synth takes an interesting approach to synthesis, offering a trio of synth engines that each op...
15/06/2026
Applications now open for 2026
The Oram Awards have returned for 2026 to celebrate the unusual, unique and unfiltered creative worlds of women and gender-di...
15/06/2026
New intelligent auto-fader plug-in revealed
PSPaudioware's latest release offers automatic level adjustment and provides more detailed control than many...
15/06/2026
4.78M AUSSIES TUNE IN FOR SOCCEROOS WIN OVER T RK YE ON SBS
15 June, 2026
Media releases
Match had a Total TV average audience of 3.035 million, with over ...
15/06/2026
SBS Head of Commissioning John Godfrey to depart after 18 years
15 June, 2026
Media releases
SBS Head of Commissioning John Godfrey will depart the broadca...
15/06/2026
Greater Manchester Police installs Rohde & Schwarz security scanner for custody ...
15/06/2026
Insights from NAGRAVISION's latest industry webinar featuring One Hungary, Liberty Global and Media Press Group
In this blog, Laura Rognoni explores the k...
15/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
15/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
15/06/2026
Clear-Com has introduced Avalon , a purpose built 1RU IP intercom communication platform for modern networked production, designed to simplify and scale workfl...
15/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
15/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
15/06/2026
MiLB Club Deploys LDX 110 Cameras at CarMax Park to Deliver A New Standard in Engaging Fan Experience
Grass Valley today announced that the Richmond Flying Sq...
15/06/2026
Detach from Direct-Attached: How Remote Editing with EVO Keeps Creative Teams Mo...
15/06/2026
Techtel Completes Media Production Setup for a major AFL sporting organisation
Sports
15 June Written By Suzanne Costello
(Sydney, Australia 15 June 2026)...
15/06/2026
Monday 15 June 2026
Sky News takes viewers inside Minab in new film investigati...
15/06/2026
Fox Corporation to Acquire Roku, Inc. Combination Creates a Scaled Media and Technology Platform with Superior Reach, Engagement and Monetization Capability
...
14/06/2026
Library captures 1960s R&B/pop drum sound
Following on from their recent wave of plug-in effects, Iconic Instruments have just launched an all-new virtual d...
14/06/2026
HBO Comedy Rooster Shot with URSA Cine 17K 65
Brie Clayton June 14, 2026
0 Comments
Large format brings viewers intimately close to characters.
Black...
13/06/2026
Latest expansion pack includes 252 presets
Devious Machines have recently introduced another expansion for their powerful multi-effects plug-in, Infiltrator...
13/06/2026
Create custom DAW/plug-in controllers using prompts
MetaGrid have recently introduced an all-new AI Builder function to their touchscreen-based control surf...
13/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
13/06/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/06/2026
YES Network and The Gotham Sports App will air seven Athletes Unlimited Softball...
12/06/2026
The United Football League will host its FAST Innovation Suite at the 2026 United Bowl presented by Credit One Bank on Saturday, June 13 at 3:00 p.m. ET at Audi...
12/06/2026
PTZOptics and LayerJot will present live demonstrations at InfoComm 2026 showing how natural-language AI prompting, robotic camera control, and on-device comput...
12/06/2026
MultiDyne Video and Fiber Optic Systems will exhibit at InfoComm 2026, featuring...
12/06/2026
Ateme has announced that Eurovision Services is using Ateme's software-based frame-rate conversion technology for international live event workflows. The de...
12/06/2026
Bitmovin and Simplestream have announced a partnership with Xperi to simplify the launch of OTT streaming services on TiVo OS smart TVs and devices. The collabo...
12/06/2026
Net Insight has announced that a multinational technology company is deploying a...
12/06/2026
MLB Players Inc., the business arm of the MLB Players Association, has announced a partnership with Athletes First to develop and sell brand partnerships across...
12/06/2026
Guntermann and Drunck (G&D) and VuWall have announced the CommandKeyboard-Advanc...
12/06/2026
Comcast Smart Solutions announces a new smart technology deployment with Major L...
12/06/2026
Elevation Worship completed the initial leg of its Elevation Nights 2026 tour ...
12/06/2026
AJA Video Systems has announced KONA IP25 support for Colorfront Transkoder and ...
12/06/2026
Audinate Group Limited (ASX: AD8) will exhibit at InfoComm 2026 (Booth C7321, Ce...
12/06/2026
Pac-12 Commissioner Teresa Gould has announced the appointment of Scott Adametz as Chief Technology Officer. The Pac-12 describes the hire as the first CTO appo...
12/06/2026
Grass Valley has announced AMPP Edge Live, a production system combining Grass Valley hardware, NVIDIA Blackwell GPU acceleration, and AMPP OS in a single platf...