
Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow without multiplying operational complexity.
NVIDIA's latest work with Amazon Web Services (AWS) addresses each of those constraints. Across Amazon OpenSearch and Amazon EC2, NVIDIA AI infrastructure is giving enterprises more practical paths to deploy AI at production scale.
EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs expand the compute layer for AI, graphics, video and data analytics workloads, while the NVIDIA cuVS library accelerates the retrieval layer by making GPU-powered vector indexing the default in OpenSearch Serverless. And with AWS achieving NVIDIA Exemplar Cloud status for NVIDIA GB300, customers can trust they're receiving peak optimized performance for their training workloads.
NVIDIA RTX PRO 4500 Blackwell Server Edition Multi-Workload GPUs Power New Amazon EC2 G7 Instances Amazon EC2 G7 instances bring NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs to AWS for AI inference, graphics, spatial computing and GPU-accelerated data analytics - delivering a new instance type engineered for production workloads that need performance without the operational overhead of a customer-managed GPU platform.
Compared with G6 instances, G7 delivers up to 4.6x AI inference performance, up to 2.1x graphics performance and significantly faster GPU-accelerated data analytics on Amazon EMR using the NVIDIA cuDF library for Apache Spark workloads.
With support for up to eight GPUs, 256GB of total GPU memory, 700 Gbps of EFA-enabled networking and up to 7.6TB of local NVMe SSD storage - across one-, two-, four- and eight- GPU configurations plus bare metal, coming soon - G7 instances let customers right-size infrastructure for their workloads instead of over-provisioning for them.
The platform's versatility means AI teams get lower-latency inference. Media and entertainment teams get high-resolution video workflows and rendering. Simulation, computer-aided design, virtual desktop infrastructure, gaming and spatial computing teams get the same instance type for graphics-intensive applications. And data teams can apply the GPU memory, local storage and networking improvements to analytics pipelines and vector database workloads.
G7 instances are accessible through AWS Deep Learning Amazon Machine Images (AMIs), Amazon Deep Learning Containers, Amazon EMR, Amazon EKS, Amazon ECS and graphics AMIs - and coming soon to Amazon SageMaker AI.
NVIDIA cuVS Makes GPU-Accelerated Vector Search the Default in Amazon OpenSearch The next generation of Amazon OpenSearch Serverless powers agentic AI and dynamic workloads with no infrastructure management required. It uses GPU-accelerated vector indexing, powered by NVIDIA cuVS, as the default compute choice for all vector collections.
For teams building retrieval-augmented generation, semantic search, recommendation systems and agentic AI applications, that shift matters. It turns GPU-powered vector search from a specialized optimization project into a standard AWS capability.
The customer impact is direct: vector indexing up to 10x faster at a quarter of the cost, compared with CPU-only builds - making billion-scale vector databases practical to build in under an hour.
By making NVIDIA cuVS the default in OpenSearch Serverless, AWS customers get a much faster path from raw data to production-ready AI retrieval infrastructure - with serverless scaling that reduces operational overhead when workloads are idle.
AWS Achieves NVIDIA Exemplar Cloud Status for GB300 Training Performance AWS has achieved NVIDIA Exemplar Cloud status on NVIDIA GB300 for training workloads. This means AWS meets the rigorous performance thresholds that NVIDIA uses to benchmark AI workloads against its reference architecture.
This achievement is the result of deep co-engineering efforts between AWS and NVIDIA teams. Through the NVIDIA Exemplar Clouds initiative, developers and AI leaders can be confident they're using consistent, high-performance cloud infrastructure for large-scale training, helping teams evaluate cloud providers with greater confidence, improve total cost of ownership and move AI projects from planning to production more efficiently.
Together, these advancements reinforce every layer of the AI infrastructure stack on AWS. The throughline is the same: production-grade AI infrastructure that performs at scale, without adding operational burden to the teams running it.
Learn more in this AWS blog.
NVIDIA GTC Berlin Registration Is Now Open October 20-22
Register Now
Recent News
AI
NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science June 30, 2026
AI Infrastructure
How NVIDIA's Inference Software Stack Powers the Lowest Token Cost June 30, 2026
Robotics
Into the Omniverse: Three Workflows for Improving Vision AI Agent Accuracy With Synthetic Data and Fine-Tuning June 30, 2026
AI Infrastructure
Claude Meets Blackwell Ultra: Anthropic's Models Now Run on NVIDIA GB300 in Azure June 29, 2026
View All Recent News
Categories:
AI Infrastructure
Cloud
Tags:
Agentic AI
NVIDIA Blackwell
More from Nvidia
06/08/2026
Editor's note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners and enterprises can transform their workflows u...
06/08/2026
August is here, bringing 26 new games for GeForce NOW members.
Command the seas in World of Warships: Legends and discover what's next in the GeForce NOW ...
04/08/2026
Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory.
But rising needs aren't me...
04/08/2026
For robotaxis and other autonomous vehicles (AVs), the hardest problems aren'...
04/08/2026
Members of the Open Secure AI Alliance - now more than 120 organizations strong - are developing new guidelines to strengthen agentic AI cybersecurity as the an...
30/07/2026
Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all.
With cloud gaming, everyday laptops used for clas...
28/07/2026
As a discerning AI investor who values style and substance, Sarah Guo knows this...
27/07/2026
Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government...
26/07/2026
The complexity of modern chip design continues to grow as engineering teams work to develop increasingly sophisticated CPUs, GPUs and AI systems. To help meet t...
23/07/2026
At this week's AI Summit in San Francisco, South Korean President Jae Myung Lee and some of the country's top business leaders and researchers are meeti...
23/07/2026
Lock in and load up the cloud. GFN Thursday brings fresh updates and new adventu...
22/07/2026
NVIDIA founder and CEO Jensen Huang today visited the Naval Postgraduate School in Monterey, California, to commission an NVIDIA DGX GB300 system - bringing one...
22/07/2026
Before a healthcare robot can be useful in the real world, it has to learn how the physical world pushes back. Anatomy varies. Instruments bend, press, slip and...
21/07/2026
The AI era runs on AI infrastructure. Many of these advanced systems are built a...
21/07/2026
AI has entered the gigascale era.
The world's most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier mod...
21/07/2026
NVIDIA Vera Rubin is here, and it's going gigascale.
Vera Rubin NVL72 produ...
20/07/2026
At this year's SIGGRAPH conference, running through Thursday, July 23, in Lo...
20/07/2026
Erin Davis calls it the SuperDuperPOD. That's two things in one name: phar...
17/07/2026
Think of a professional athlete. What separates elite performers is what happens...
16/07/2026
Onimusha: Way of the Sword is coming to GeForce NOW at launch, with the playable...
15/07/2026
General-purpose robots and autonomous machines are moving from research labs to ...
15/07/2026
Home to leading manufacturers, robotics pioneers, infrastructure builders and iconic gaming companies, of course, Japan is one of the world's centers of AI ...
14/07/2026
Editor's note: This post is part of the Nemotron Labs blog series, which exp...
14/07/2026
Power is AI infrastructure's inescapable constraint. How many tokens an AI factory can generate within a fixed power budget determines its revenue and profi...
09/07/2026
This GFN Thursday brings more games, more power and more ways to play on GeForce NOW.
The cloud gaming service is expanding with a new GeForce RTX 5080-powere...
08/07/2026
NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top c...
07/07/2026
Max single-threaded CPUs at scale are a new category of CPUs built for the agentic AI era.
Across the creation and deployment of an agentic system, the CPU is...
06/07/2026
Open source AI has shown how quickly developers can innovate when models, data a...
06/07/2026
Nations have long invested in domestic infrastructure to advance their economies, protect and use their data, and take advantage of technology opportunities in ...
06/07/2026
Every year, the International Conference on Machine Learning (ICML) reveals where thousands of AI researchers have decided to put their work.
This year's ...
02/07/2026
Summer is heating up - and GeForce NOW is taking players along for the ride.
Start the month with Monopoly: Star Wars Heroes vs. Villains, bringing a galaxy fa...
01/07/2026
As AI moves from model development to production inference, compute demand is ac...
30/06/2026
Life sciences has entered an era of computational scale, and for more than a dec...
30/06/2026
As organizations move from AI pilots to production AI factories, infrastructure decisions have shifted from peak chip specifications to cost per token: how many...
30/06/2026
Editor's note: This post is part of Into the Omniverse, a series focused on ...
29/06/2026
Anthropic's Claude models in Microsoft Foundry - hosted on Microsoft Azure a...
29/06/2026
Showcasing the importance of open source innovation in American AI, Palantir'...
25/06/2026
Summer savings are heating up. From the Steam Summer Sale to GeForce NOW membership discounts, this week's GFN Thursday delivers double the deals and more w...
23/06/2026
Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow wi...
23/06/2026
News Highlights:
NVIDIA technology runs 81% of the TOP500 and 90% of the systems new to the list.
26 systems on the TOP500 adopted the NVIDIA Grace CPU, up ei...
23/06/2026
Editor's note: This post is part of the Nemotron Labs blog series, which explores how the latest open models, datasets and training techniques help business...
22/06/2026
Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that i...
22/06/2026
The next era of AI will not be defined by compute alone. Its growth will be dete...
22/06/2026
Mission, Vision and Veritas - new Los Alamos National Laboratory (LANL) supercom...
22/06/2026
At the ISC conference running in Hamburg this week, NVIDIA is introducing new so...
22/06/2026
For the past two years, the U.S. National Science Foundation's National Arti...
22/06/2026
JUPITER, Europe's first exascale supercomputer at Germany's Forschungszentrum J lich, runs on NVIDIA Grace Hopper Superchips and NVIDIA Quantum-X800 Inf...
21/06/2026
Hot tubs sit at about 38 to 40 degrees Celsius, warm enough that most people can only soak for about 15 minutes. NVIDIA's newest AI servers can run their co...
18/06/2026
In a consequential grid infrastructure decision, the Federal Energy Regulatory C...
18/06/2026
Play favorite titles from popular game libraries, keep progress synced and jump ...