
For the past 500 years, the National Library of Sweden has collected virtually every word published in Swedish, from priceless medieval manuscripts to present-day pizza menus.
Thanks to a centuries-old law that requires a copy of everything published in Swedish to be submitted to the library - also known as Kungliga biblioteket, or KB - its collections span from the obvious to the obscure: books, newspapers, radio and TV broadcasts, internet content, Ph.D. dissertations, postcards, menus and video games. It's a wildly diverse collection of nearly 26 petabytes of data, ideal for training state-of-the-art AI.
We can build state-of-the-art AI models for the Swedish language since we have the best data, said Love B rjeson, director of KBLab, the library's data lab.
Using NVIDIA DGX systems, the group has developed more than two dozen open-source transformer models, available on Hugging Face. The models, downloaded by up to 200,000 developers per month, enable research at the library and other academic institutions.
Before our lab was created, researchers couldn't access a dataset at the library - they'd have to look at a single object at a time, B rjeson said. There was a need for the library to create datasets that enabled researchers to conduct quantity-oriented research.
With this, researchers will soon be able to create hyper-specialized datasets - for example, pulling up every Swedish postcard that depicts a church, every text written in a particular style or every mention of a historical figure across books, newspaper articles and TV broadcasts.
Turning Library Archives Into AI Training Data The library's datasets represent the full diversity of the Swedish language - including its formal and informal variations, regional dialects and changes over time.
Our inflow is continuous and growing - every month, we see more than 50 terabytes of new data, said B rjeson. Between the exponential growth of digital data and ongoing work digitizing physical collections that date back hundreds of years, we'll never be finished adding to our collections.
The library's archives include audio, text and video. Soon after KBLab was established in 2019, B rjeson saw the potential for training transformer language models on the library's vast archives. He was inspired by an early, multilingual, natural language processing model by Google that included 5GB of Swedish text.
KBLab's first model used 4x as much - and the team now aims to train its models on at least a terabyte of Swedish text. The lab began experimenting by adding Dutch, German and Norwegian content to its datasets after finding that a multilingual dataset may improve the AI's performance.
NVIDIA AI, GPUs Accelerate Model Development The lab started out using consumer-grade NVIDIA GPUs, but B rjeson soon discovered his team needed data-center-scale compute to train larger models.
We realized we can't keep up if we try to do this on small workstations, said B rjeson. It was a no-brainer to go for NVIDIA DGX. There's a lot we wouldn't be able to do at all without the DGX systems.
The lab has two NVIDIA DGX systems from Swedish provider AddPro for on-premises AI development. The systems are used to handle sensitive data, conduct large-scale experiments and fine-tune models. They're also used to prepare for even larger runs on massive, GPU-based supercomputers across the European Union - including the MeluXina system in Luxembourg.
Our work on the DGX systems is critically important, because once we're in a high-performance computing environment, we want to hit the ground running, said B rjeson. We have to use the supercomputer to its fullest extent.
The team has also adopted NVIDIA NeMo Megatron, a PyTorch-based framework for training large language models, with NVIDIA CUDA and the NVIDIA NCCL library under the hood to optimize GPU usage in multi-node systems.
We rely to a large extent on the NVIDIA frameworks, B rjeson said. It's one of the big advantages of NVIDIA for us, as a small lab that doesn't have 50 engineers available to optimize AI training for every project.
Harnessing Multimodal Data for Humanities Research In addition to transformer models that understand Swedish text, KBLab has an AI tool that transcribes sound to text, enabling the library to transcribe its vast collection of radio broadcasts so that researchers can search the audio records for specific content.
AI-enhanced databases are the latest evolution of library records, which were long stored in physical card catalogs. KBLab is also starting to develop generative text models and is working on an AI model that could process videos and create automatic descriptions of their content.
We also want to link all the different modalities, B rjeson said. When you search the library's databases for a specific term, we should be able to return results that include text, audio and video.
KBLab has partnered with researchers at the University of Gothenburg, who are developing downstream apps using the lab's models to conduct linguistic research - including a project supporting the Swedish Academy's work to modernize its data-driven techniques for creating Swedish dictionaries.
The societal benefits of these models are much larger than we initially expected, B rjeson said.
Images courtesy of Kungliga biblioteket
Most recent headlines
05/01/2027
Worlds first 802.15.4ab-UWB chip verified by Calterah and Rohde & Schwarz to be ...
01/06/2026
January 6 2026, 05:30 (PST) Dolby Sets the New Standard for Premium Entertainment at CES 2026
Throughout the week, Dolby brings to life the latest innovatio...
02/05/2026
Dalet, a leading technology and service provider for media-rich organizations, t...
01/05/2026
January 5 2026, 18:30 (PST) NBCUniversal's Peacock to Be First Streamer to ...
01/04/2026
January 4 2026, 18:00 (PST) DOLBY AND DOUYIN EMPOWER THE NEXT GENERATON OF CREATORS WITH DOLBY VISION
Douyin Users Can Now Create And Share Videos With Stun...
12/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
12/02/2026
The GeForce NOW sixth-anniversary festivities roll on this February, continuing a monthlong celebration of NVIDIA's cloud gaming service.
This week brings ...
12/02/2026
TIME100 Health list features Scripps Research Professor Darrell Irvine Irvine is recognized for his work in empowering the immune system to fight disease, which...
11/02/2026
FYI: Phone Support Maintenance One thing we pride ourselves on here at Utah Scientific is our 24-hour support included with our signature 10-year hardware warra...
11/02/2026
Leading provider of video streaming solutions, Bitmovin, has appointed Ian Baglow as Co-CEO alongside existing CEO and Co-Founder Stefan Lederer. Under this str...
11/02/2026
Paramount and the CBS Television Network will partner to air UFC 326: HOLLOWAY vs. OLIVEIRA 2 live on Saturday, March 7, from T-Mobile Arena in Las Vegas, mar...
11/02/2026
Beginning February 10, fans can buy MLB.TV on ESPN, a new milestone in one of sports media's longest-standing partnerships. ESPN becomes the new streaming h...
11/02/2026
Fubo Sports Network is available to Hulu's Live TV subscribers in the core $89.99 a month subscription plan, which also includes full access to the entire H...
11/02/2026
Following a competitive public tender process, Rai (Radiotelevisione Italiana), the national public broadcasting company of Italy, has awarded Imagine Communica...
11/02/2026
Major League Baseball is making in-market streaming subscriptions for 20 Clubs available today for fans. Subscriptions for the following Clubs are available vi...
11/02/2026
Building on successful demonstrations during the Paris Olympics 2024, Italian public service broadcaster Rai and the European Broadcasting Union (EBU) are condu...
11/02/2026
Following Sunday's Super Bowl LX, ESPN and Disney unveiled We're Going,...
11/02/2026
Delayed streams are a growing source of frustration for sports fans. During the 2026 Super Bowl, some streams lagged up to 62 seconds behind the action on the f...
11/02/2026
NASCAR and FloSports announces an expanded slate of racing events that will bring FloRacing coverage live throughout the 2026 season to the NASCAR Channel, furt...
11/02/2026
Manifold technologies GmbH announces the appointment of Nick Tucker as Sales Manager for Europe, reinforcing the company's continued growth across broadcast...
11/02/2026
Genies, the AI avatar technology company powering the next era of interactive digital identity, entered into a landmark collaboration with MLB Players, Inc., th...
11/02/2026
The International Cricket Council (ICC) and Google have joined forces for an AI-...
11/02/2026
Dolby's CEO Kevin Yeaman and Giles Baker, SVP of Dolby Cloud Solutions, shared how the brand's latest innovations - Dolby Vision, Dolby Atmos, and Dolby...
11/02/2026
Ilitch Sports + Entertainment has entered a first of its kind partnership with Major League Baseball, which will provide broadcast support to both the Detroit T...
11/02/2026
For major U.S. events like Super Bowl 2026, FIFA World Cup, America 250, and the...
11/02/2026
Broadcasts of the NHL's Detroit Red Wings will also be produced by the leagu...
11/02/2026
Video moves fast can your DAM keep up?
Join Blue Lucy in LA for the West Coast's leading Digital Asset Management event as we explore, celebrate, and acc...
11/02/2026
NEW YORK - February 10, 2026 - An estimated 124.9 million viewers watched Super Bowl LX on Sunday, February 8, according to Nielsen's Big Data Panel measu...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Clear-Com provided an advanced, IP-based communications infrastructure for TEDNext 2025, supporting production, media, and editorial teams with a highly flexib...
11/02/2026
Astera introduces QuikBeam, the newest addition to its acclaimed Quik family of focusing LED Fresnels. This ultra-compact spotlight combines the equivalent powe...
11/02/2026
Following a competitive public tender process, Rai (Radiotelevisione Italiana), the national public broadcasting company of Italy, has awarded Imagine Communica...
11/02/2026
With Convertible Mount for NL Bowens & Aputure A Mounts See it at BSC Expo Stand #133 LCA
DoPchoice continues to refine light shaping tools for professional LE...
11/02/2026
World Premiere at BSC Expo, Booth #319 Oberkochen/Germany, 10 February 2026
ZEISS introduces the new Aatma, set of nine high-end full frame T1.5 cinema primes ...
11/02/2026
As Re-recording Mixer and Head of Sound at The Farm, one of UK's leading post-production facilities, Nick Fry has built his career on making stories sound a...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Share
Copy link
Facebook
X
Linkedin
Bluesky
Email...
11/02/2026
Graduate Spotlight: Gabrielle Rodriguez The educator, who grew up in the Philippines, shares how shes bringing what she learned at Berklee back home.
Februar...
11/02/2026
Wednesday 11 February 2026
Sky brings together Netflix, Disney , HBO Max and Ha...
11/02/2026
Back to All News
Netflix Confirms Production of Love O'Clock' From the...
11/02/2026
Back to All News
Investing in Belgian Stories: A Commitment to Culture and Choice
From left to right: Undercover, Ang le, Rough Diamonds, Into the Night, John...
11/02/2026
At the end of January, ICG headed off to the Portuguese capital, Lisbon, for our annual conference.
An early flight gave us plenty of time to start exploring s...