In a landmark development for the global cloud computing and semiconductor industries, Amazon and NVIDIA have officially announced a massive expansion of their long-standing partnership, signaling a transformative shift in the infrastructure powering the next generation of artificial intelligence. Following the release of NVIDIA’s fiscal second-quarter earnings, the companies confirmed a strategic agreement that will see Amazon Web Services (AWS) deploy up to three million NVIDIA AI GPUs across its global infrastructure. This expansion is designed to meet the skyrocketing demand for agentic AI, physical AI, and industrial automation, marking one of the largest hardware commitments in the history of cloud computing.
The deal represents a significant escalation from previous agreements. During NVIDIA’s GTC conference in March, the two companies had outlined a plan for the deployment of one million GPUs. The latest announcement triples that figure, adding two million additional units to the AWS ecosystem. This surge in capacity is aimed at providing the computational "horsepower" necessary to sustain the increasingly complex workloads of generative AI and autonomous systems, which require unprecedented levels of parallel processing and memory bandwidth.
The Integration of NVIDIA Vera CPUs and the Rise of Agentic AI
A pivotal technical highlight of this expanded collaboration is the introduction of NVIDIA’s Vera CPUs into the AWS platform. While GPUs have traditionally been the primary focus of AI infrastructure, the architectural requirements of "agentic AI"—systems capable of independent reasoning, planning, and multi-step task execution—have shifted the spotlight back toward high-performance central processing.
As of 2026, the industry has observed that the role of the CPU in AI clusters is no longer merely administrative. Agentic AI requires sophisticated logic handling and coordination that benefits from the tight integration of specialized CPUs with GPU clusters. The Vera CPU, part of NVIDIA’s latest "Vera Rubin" architecture, is engineered to handle these complex orchestration tasks. By incorporating Vera CPUs into its instances, AWS aims to offer a more balanced and efficient compute environment.
Data released alongside the announcement highlights the performance leap offered by this new hardware. The Vera Rubin NVL72 platform is reported to deliver up to 30 times better throughput per megawatt compared to the previous generation GB300 NVL72 systems. This efficiency is critical for cloud providers like Amazon, who are facing increasing pressure to manage the massive energy consumption associated with hyperscale AI data centers.
Custom Silicon and NVHBM Memory Innovation
While the deployment of three million NVIDIA GPUs is a cornerstone of the deal, Amazon continues to pursue a "broadest possible" compute strategy. This involves a hybrid approach where NVIDIA’s hardware complements Amazon’s in-house custom silicon, specifically the Trainium chips designed for deep learning training.

To enhance the performance of its proprietary hardware, Amazon will leverage NVIDIA’s newly announced NVHBM (NVIDIA High Bandwidth Memory) technology. This custom memory solution is designed to bridge the gap between standard HBM4E and the specific needs of high-end AI accelerators. According to technical specifications, NVHBM provides a 30% increase in bandwidth and 15% higher power efficiency over HBM4E memory. By integrating NVHBM with Trainium chips, AWS can offer its customers a high-performance alternative to standard GPU instances, potentially lowering costs for specific training workloads while maintaining top-tier memory speeds.
This synergy between NVIDIA’s memory technology and Amazon’s chip design illustrates a deeper level of technical integration than seen in previous years. It suggests that the relationship between the two firms has evolved from a simple vendor-customer dynamic into a co-engineering partnership aimed at optimizing the entire hardware stack.
Sovereign AI and the US Government AI Factories
A significant portion of the new GPU deployment is earmarked for the public sector. Amazon revealed that 100,000 NVIDIA GPUs will be dedicated to building "AI factories" specifically for the United States government. These facilities are designed to meet the stringent security requirements of Impact Level 6 (IL6) classifications.
IL6 is the security standard required for processing and storing information categorized as "Secret" by the Department of Defense and other federal agencies. By establishing these dedicated AI factories, Amazon and NVIDIA are positioning themselves as the primary infrastructure providers for "Sovereign AI"—the concept of a nation-state maintaining its own AI capabilities and data security within its borders. These government-specific clusters will likely be used for national security simulations, advanced cryptography, and autonomous defense systems, ensuring that the U.S. government has access to the same cutting-edge hardware as the private sector while maintaining strict data isolation.
A Chronology of the AWS-NVIDIA Partnership
The relationship between AWS and NVIDIA spans 16 years, tracing the evolution of the cloud from basic web hosting to the epicenter of the AI revolution.
- 2010: AWS launched the first GPU-accelerated cloud instances, allowing researchers to run parallel workloads without owning physical hardware.
- 2023: As the generative AI boom took hold, AWS became one of the first major providers to offer NVIDIA H100 Tensor Core GPUs at scale.
- March 2024: At GTC, the companies announced the "Project Ceiba" supercomputer, a collaboration to build one of the world’s fastest AI supercomputers on AWS, initially targeting a deployment of one million GPUs.
- Late 2024 – 2025: The partnership expanded to include Blackwell-architecture GPUs and the integration of NVIDIA’s NIM (NVIDIA Inference Microservices) on AWS.
- 2026 (Current): The expansion to three million GPUs and the introduction of the Vera Rubin architecture marks the "full stack" era of the partnership, encompassing CPUs, networking, and custom memory.
Industry Perspectives and Leadership Vision
The scale of this agreement reflects a market reality where demand for AI compute continues to outpace supply. Jensen Huang, founder and CEO of NVIDIA, emphasized that the partnership is moving beyond simple hardware provisioning.
"NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast," Huang stated. "For 16 years, we have scaled NVIDIA computing in the cloud together. Now we are expanding our partnership across the full stack—GPUs, CPUs, networking, open models, and software—to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver."

Industry analysts suggest that this deal is a defensive and offensive move for both companies. For Amazon, it ensures that it remains the "first choice" for enterprise AI, preventing customers from migrating to Microsoft Azure or Google Cloud due to hardware shortages. For NVIDIA, it secures a massive, long-term revenue stream and cements its Vera CPUs and NVHBM memory as industry standards.
Broader Economic and Technological Implications
The deployment of three million GPUs has profound implications for the global supply chain and the trajectory of AI development. The sheer volume of hardware involved will likely consume a significant portion of NVIDIA’s production capacity for the coming fiscal years, potentially impacting the availability of chips for smaller players.
Furthermore, the focus on "Physical AI" mentioned by Jensen Huang points toward a future where AI is no longer confined to chatbots and image generators. Physical AI refers to the application of AI in the real world—robotics, autonomous manufacturing, and self-driving logistics. By providing the backend infrastructure for these applications, AWS and NVIDIA are laying the groundwork for the next industrial revolution, where software-defined intelligence controls physical hardware at scale.
The environmental impact also remains a topic of analysis. With the Vera Rubin architecture delivering 30x better throughput per megawatt, the companies are attempting to decouple computational growth from energy growth. However, the sheer scale of three million GPUs will still require massive investments in power grid infrastructure and cooling technologies, areas where Amazon is already investing heavily through its commitment to carbon neutrality and nuclear energy projects.
As the fiscal year progresses, the tech industry will be watching closely to see how quickly these three million units can be brought online. This collaboration effectively raises the "barrier to entry" for cloud competition, as the capital expenditure required to match such a fleet of NVIDIA hardware is now measured in the tens of billions of dollars. For now, the Amazon-NVIDIA alliance stands as the most formidable infrastructure partnership in the digital age, dictating the pace of innovation for the global AI economy.







