
Enterprises are rapidly adopting generative AI, large language models (LLMs), advanced graphics and digital twins to increase operational efficiencies, reduce costs and drive innovation.
However, to adopt these technologies effectively, enterprises need access to state-of-the-art, full-stack accelerated computing platforms. To meet this demand, Oracle Cloud Infrastructure (OCI) today announced NVIDIA L40S GPU bare-metal instances available to order and the upcoming availability of a new virtual machine accelerated by a single NVIDIA H100 Tensor Core GPU. This new VM expands OCI's existing H100 portfolio, which includes an NVIDIA HGX H100 8-GPU bare-metal instance.
Paired with NVIDIA networking and running the NVIDIA software stack, these platforms deliver powerful performance and efficiency, enabling enterprises to advance generative AI.
NVIDIA L40S Now Available to Order on OCI The NVIDIA L40S is a universal data center GPU designed to deliver breakthrough multi-workload acceleration for generative AI, graphics and video applications. Equipped with fourth-generation Tensor Cores and support for the FP8 data format, the L40S GPU excels in training and fine-tuning small- to mid-size LLMs and in inference across a wide range of generative AI use cases.
For example, a single L40S GPU (FP8) can generate up to 1.4x more tokens per second than a single NVIDIA A100 Tensor Core GPU (FP16) for Llama 3 8B with NVIDIA TensorRT-LLM at an input and output sequence length of 128.
The L40S GPU also has best-in-class graphics and media acceleration. Its third-generation NVIDIA Ray Tracing Cores (RT Cores) and multiple encode/decode engines make it ideal for advanced visualization and digital twin applications.
The L40S GPU delivers up to 3.8x the real-time ray-tracing performance of its predecessor, and supports NVIDIA DLSS 3 for faster rendering and smoother frame rates. This makes the GPU ideal for developing applications on the NVIDIA Omniverse platform, enabling real-time, photorealistic 3D simulations and AI-enabled digital twins. With Omniverse on the L40S GPU, enterprises can develop advanced 3D applications and workflows for industrial digitalization that will allow them to design, simulate and optimize products, processes and facilities in real time before going into production.
OCI will offer the L40S GPU in its BM.GPU.L40S.4 bare-metal compute shape, featuring four NVIDIA L40S GPUs, each with 48GB of GDDR6 memory. This shape includes local NVMe drives with 7.38TB capacity, 4th Generation Intel Xeon CPUs with 112 cores and 1TB of system memory.
These shapes eliminate the overhead of any virtualization for high-throughput and latency-sensitive AI or machine learning workloads with OCI's bare-metal compute architecture. The accelerated compute shape features the NVIDIA BlueField-3 DPU for improved server efficiency, offloading data center tasks from CPUs to accelerate networking, storage and security workloads. The use of BlueField-3 DPUs furthers OCI's strategy of off-box virtualization across its entire fleet.
OCI Supercluster with NVIDIA L40S enables ultra-high performance with 800Gbps of internode bandwidth and low latency for up to 3,840 GPUs. OCI's cluster network uses NVIDIA ConnectX-7 NICs over RoCE v2 to support high-throughput and latency-sensitive workloads, including AI training.
We chose OCI AI infrastructure with bare-metal instances and NVIDIA L40S GPUs for 30% more efficient video encoding, said Sharon Carmel, CEO of Beamr Cloud. Videos processed with Beamr Cloud on OCI will have up to 50% reduced storage and network bandwidth consumption, speeding up file transfers by 2x and increasing productivity for end users. Beamr will provide OCI customers video AI workflows, preparing them for the future of video.
Single-GPU H100 VMs Coming Soon on OCI The VM.GPU.H100.1 compute virtual machine shape, accelerated by a single NVIDIA H100 Tensor Core GPU, is coming soon to OCI. This will provide cost-effective, on-demand access for enterprises looking to use the power of NVIDIA H100 GPUs for their generative AI and HPC workloads.
A single H100 provides a good platform for smaller workloads and LLM inference. For example, one H100 GPU can generate more than 27,000 tokens per second for Llama 3 8B (up to 4x more throughput than a single A100 GPU at FP16 precision) with NVIDIA TensorRT-LLM at an input and output sequence length of 128 and FP8 precision.
The VM.GPU.H100.1 shape includes 2 3.4TB of NVMe drive capacity, 13 cores of 4th Gen Intel Xeon processors and 246GB of system memory, making it well-suited for a range of AI tasks.
Oracle Cloud's bare-metal compute with NVIDIA H100 and A100 GPUs, low-latency Supercluster and high-performance storage delivers up to 20% better price-performance for Altair's computational fluid dynamics and structural mechanics solvers, said Yeshwant Mummaneni, chief engineer of data management analytics at Altair. We look forward to leveraging these GPUs with virtual machines for the Altair Unlimited virtual appliance.
GH200 Bare-Metal Instances Available for Validation OCI has also made available the BM.GPU.GH200 compute shape for customer testing. It features the NVIDIA Grace Hopper Superchip and NVLink-C2C, a high-bandwidth, cache-coherent 900GB/s connection between the NVIDIA Grace CPU and NVIDIA Hopper GPU. This provides over 600GB of accessible memory, enabling up to 10x higher performance for applications running terabytes of data compared to the NVIDIA A100 GPU.
Optimized Software for Enterprise AI Enterprises have a wide variety of NVIDIA GPUs to accelerate their AI, HPC and data analytics workloads on OCI. However, maximizing the full potential of these GPU-accelerated compute instances requires an optimized software layer.
NVIDIA NIM, part of the NVIDIA AI Enterprise software platform available on the OC
More from Nvidia
15/07/2025
Submissions for NVIDIA's Plug and Play: Project G-Assist Plug-In Hackathon a...
14/07/2025
This month, NVIDIA founder and CEO Jensen Huang promoted AI in both Washington, D.C. and Beijing - emphasizing the benefits that AI will bring to business and s...
11/07/2025
Ceramics - the humble mix of earth, fire and artistry - have been part of a global conversation for millennia.
From Tang Dynasty trade routes to Renaissance pa...
10/07/2025
In the race to understand our planet's changing climate, speed and accuracy are everything. But today's most widely used climate simulators often strugg...
10/07/2025
As one of the world's largest emerging markets, Indonesia is making strides toward its Golden 2045 Vision - an initiative tapping digital technologies and...
10/07/2025
Grab a friend and climb toward the clouds - PEAK is now available on GeForce NOW, enabling members to try the hugely popular indie hit on virtually any device.
...
10/07/2025
Coding assistants or copilots - AI-powered assistants that can suggest, explain and debug code - are fundamentally changing how software is developed for both e...
08/07/2025
Modern AI applications increasingly rely on models that combine huge parameter c...
03/07/2025
The forecast this month is showing a 100% chance of epic gaming. Catch the scorching lineup of 20 titles coming to the cloud, which gamers can play whether indo...
02/07/2025
Black Forest Labs, one of the world's leading AI research labs, just changed the game for image generation.
The lab's FLUX.1 image models have earned g...
01/07/2025
In many parts of the world, including major technology hubs in the U.S., there's a yearslong wait for AI factories to come online, pending the buildout of n...
26/06/2025
As of today, NVIDIA now supports the general availability of Gemma 3n on NVIDIA RTX and Jetson. Gemma, previewed by Google DeepMind at Google I/O last month, in...
26/06/2025
Editor's note: This blog is a part of Into the Omniverse, a series focused o...
26/06/2025
Mark Theriault founded the startup FITY envisioning a line of clever cooling products: cold drink holders that come with freezable pucks to keep beverages cold ...
26/06/2025
This GFN Thursday rolls out a new reward and games for GeForce NOW members. Whether hunting for hot new releases or rediscovering timeless classics, members can...
24/06/2025
To get the most out of AI, optimizations are critical. When developers think about optimizing AI models for inference, model compression techniques-such as quan...
24/06/2025
To speed up AI adoption across industries, HPE and NVIDIA today launched new AI factory offerings at HPE Discover in Las Vegas.
The new lineup includes everyth...
24/06/2025
From the heart of Germany's automotive sector to manufacturing hubs across F...
19/06/2025
GeForce NOW is throwing open the vault doors to welcome the legendary Borderland series to the cloud.
Whether a seasoned Vault Hunter or new to the mayhem of P...
18/06/2025
Project G-Assist - available through the NVIDIA App - is an experimental AI assistant that helps tune, control and optimize NVIDIA GeForce RTX systems.
NVIDIA&...
17/06/2025
As a global labor shortage leaves 50 million positions unfilled across industrie...
13/06/2025
Industrial AI isn't slowing down. Germany is ready.
Following London Tech Week and GTC Paris at VivaTech, NVIDIA founder and CEO Jensen Huang's Europea...
12/06/2025
Generative AI has reshaped how people create, imagine and interact with digital ...
12/06/2025
Level up GeForce NOW experiences this summer with 40% off Performance Day Passes. Enjoy 24 hours of premium cloud gaming with RTX ON, delivering low latency and...
11/06/2025
NVIDIA is launching a comprehensive, industry-defining autonomous vehicle (AV) software platform to accelerate large-scale deployment of safe, intelligent trans...
11/06/2025
NVIDIA Research has developed an AI light switch for videos that can turn daytim...
11/06/2025
Using NVIDIA platforms, tools and libraries, European telecommunications institu...
11/06/2025
NVIDIA was today named an Autonomous Grand Challenge winner at the Computer Visi...
11/06/2025
In the face of growing labor shortages and need for sustainability, European man...
11/06/2025
AI is packing and shipping efficiency for the retail and consumer packaged goods (CPG) industries, with a majority of surveyed companies in the space reporting ...
11/06/2025
Urban populations are expected to double by 2050, which means around 2.5 billion...
11/06/2025
Telecom companies last year spent nearly $295 billion in capital expenditures an...
11/06/2025
In a new effort to advance sovereign AI for European public service media, NVIDI...
11/06/2025
At GTC Paris - held alongside VivaTech, Europe's largest tech event - NVIDIA founder and CEO Jensen Huang delivered a clear message: Europe isn't just a...
10/06/2025
Germany's Leibniz Supercomputing Centre, LRZ, is gaining a new supercomputer...
10/06/2025
With a more detailed simulation of the Earth's climate, scientists and resea...
10/06/2025
Cisco and NVIDIA are helping set a new standard for secure, scalable and high-performance enterprise AI.
Announced today at the Cisco Live conference in San Di...
09/06/2025
AI isn't waiting. And this week, neither is Europe.
At London's Olympia, under a ceiling of steel beams and enveloped by the thrum of startup pitches, ...
08/06/2025
U.K. Prime Minister Keir Starmer's ambition for Britain to be an AI maker, not an AI taker, is becoming a reality at London Tech Week.
With NVIDIA's ...
05/06/2025
GeForce NOW is a gamer's ticket to an unforgettable summer of gaming. With 25 titles coming this month and endless ways to play, the summer is going to be e...
04/06/2025
NVIDIA is working with companies worldwide to build out AI factories - speeding ...
04/06/2025
Humans learn the norms, values and behaviors of society from each other - and Bernt B rnich, founder and CEO of 1X Technologies, thinks robots should learn like...
04/06/2025
4:2:2 cameras - capable of capturing double the color information compared with most standard cameras - are becoming widely available for consumers. At the same...
02/06/2025
Editor's note: This blog, originally published on October 28, 2024, has been...
02/06/2025
Since a 7.8-magnitude earthquake hit Syria and T rkiye two years ago - leaving 5...
29/05/2025
Ready for a front-row seat to the next scientific revolution?
That's the idea behind Doudna - a groundbreaking supercomputer announced today at Lawrence Be...
29/05/2025
Large language models (LLMs), trained on datasets with billions of tokens, can generate high-quality content. They're the backbone for many of the most popu...
29/05/2025
GeForce NOW is supercharging Valve's Steam Deck with a new native app - delivering the high-quality GeForce RTX-powered gameplay members are used to on a po...
28/05/2025
Building effective agentic AI systems requires rethinking how technology interac...
27/05/2025
Over a century ago, Henry Ford pioneered the mass production of cars and engines...