Menu

Uncategorized

AWS and NVIDIA Plan 2 Million More GPUs: What Blackwell, Rubin and Rubin Ultra Mean for AI Infrastructure Buyers

NVIDIA Rubin GPU

Amazon Web Services and NVIDIA have announced one of the largest expansions of accelerated AI infrastructure yet: AWS plans to deploy 2 million additional NVIDIA GPUs across its global infrastructure during 2027 and 2028. The NVIDIA Rubin GPU is becoming a central part of AWS’s next generation of AI infrastructure.

The planned deployment will span NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs and extends far beyond accelerators alone. AWS and NVIDIA are also deepening collaboration across networking, CPUs, AI factories, open models, data processing and physical AI.

For organizations planning an AI cluster, GPU server deployment or data-center expansion, the announcement carries an important message: the next generation of AI infrastructure will increasingly be purchased and designed as a complete system, not as a collection of standalone GPUs.

The NVIDIA Rubin GPU may be the headline technology, but networking bandwidth, power density, liquid cooling, CPU architecture, rack design and software compatibility are becoming equally important procurement decisions.

NVIDIA Rubin GPU

What AWS and NVIDIA Actually Announced

AWS and NVIDIA announced the expanded collaboration on August 26, 2026.

According to the companies, the new program includes:

  • 2 million additional NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs planned for AWS infrastructure in 2027–2028
  • NVIDIA Vera CPU-based infrastructure coming to AWS
  • Expanded NVIDIA Spectrum networking collaboration
  • Integration between AWS Trainium and NVIDIA NVLink Fusion technologies
  • AI factories for U.S. government workloads, including plans involving 100,000 GPUs on secure AWS infrastructure
  • Continued NVIDIA Nemotron model availability through Amazon Bedrock and Amazon SageMaker
  • Expanded use of NVIDIA technologies for Amazon Robotics and physical AI

This follows AWS’s earlier announcement at NVIDIA GTC 2026 that it intended to add more than 1 million NVIDIA GPUs beginning in 2026.

The companies now say demand has exceeded those earlier expectations.

That matters because AWS is not simply adding another generation of cloud instances. It is preparing infrastructure that spans several NVIDIA architecture cycles simultaneously.

Why 2 Million Additional NVIDIA GPUs Matters

AI infrastructure has historically been discussed in terms of individual accelerator performance: how much memory a GPU has, how many operations it can execute, or how quickly it can train a particular model.

At hyperscale, those measurements are no longer enough.

A cluster containing tens of thousands of accelerators must move enormous amounts of data between GPUs, CPUs, memory and storage. A poorly designed network can leave expensive accelerators underutilized. Insufficient power or cooling can limit rack density. Weak orchestration can reduce effective cluster performance even when the individual GPUs are extremely fast.

AWS’s expansion therefore illustrates a broader transition from the GPU server era to the AI factory era.

The relevant purchasing unit is increasingly becoming the rack, cluster or complete data-center architecture.

For enterprise buyers, this means GPU procurement needs to be evaluated together with:

Compute capacity: The accelerator architecture and amount of GPU memory required for the workload.

Networking: High-bandwidth, low-latency connectivity between GPUs and racks.

CPU capacity: Sufficient host processing for data preparation, orchestration, agent execution and other non-GPU workloads.

Power: The electrical capacity required at rack and facility level.

Cooling: Increasingly important as accelerator and rack power density rises.

Storage and data pipelines: Fast enough to prevent expensive compute infrastructure from waiting for data.

Software: Framework, container, driver and orchestration compatibility across the entire environment.

The 2-million-GPU announcement is therefore as significant for networking and data-center infrastructure suppliers as it is for GPU manufacturers.

Blackwell Ultra vs Rubin vs Rubin Ultra

AWS plans to deploy GPUs spanning three important stages of NVIDIA’s data-center roadmap.

PlatformPosition in NVIDIA RoadmapBest Viewed AsBuyer Consideration
Blackwell UltraCurrent Blackwell-generation high-end AI platformNearer-term production AI infrastructureMost relevant when deployment cannot wait for later architecture cycles
RubinNext-generation architecture built around the Vera Rubin platformLarge-scale agentic AI, reasoning, training and inferenceRequires planning around a newer rack, CPU, networking and cooling ecosystem
Rubin UltraHigher-density successor within the Rubin generationFuture extreme-scale AI factoriesInfrastructure readiness becomes especially important because of rack density and scale

This comparison highlights an important purchasing principle: the newest GPU is not automatically the correct GPU for every organization.

Deployment date, software readiness, infrastructure requirements, budget, expected utilization and workload economics can matter more than choosing the newest architecture available.

What Is the NVIDIA Rubin GPU and Vera Rubin Platform?

NVIDIA Vera Rubin is not simply a new GPU.

It is a full AI infrastructure platform designed around multiple tightly integrated technologies, including the Rubin GPU, Vera CPU, NVLink 6, ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-6 Ethernet technologies.

A Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs, alongside high-speed networking and data-processing components.

NVIDIA positions the architecture for AI workloads involving advanced reasoning, long context windows, agentic systems, scientific computing and large-scale model training and inference.

The architectural shift is important for buyers because performance increasingly depends on communication between the different components.

Buying powerful accelerators while under-sizing the network, storage or CPU layer can produce an expensive system that never reaches its intended utilization.

Why Networking Is Becoming as Important as the GPU

Large AI models distribute computation across many accelerators.

Every time those accelerators exchange parameters, activations or training data, the network becomes part of the compute system.

This is why AWS and NVIDIA specifically included networking in their expanded partnership.

NVIDIA’s current Vera Rubin architecture incorporates technologies such as ConnectX-9 SuperNICs, Spectrum-X Ethernet and NVLink 6 for scale-up and scale-out communication.

For procurement teams, this creates a new question:

Are you purchasing GPUs, or are you purchasing a GPU cluster that can actually keep those GPUs busy?

A high-end accelerator sitting idle while waiting for data is still an idle accelerator.

Organizations planning multi-GPU servers or rack-scale AI systems should therefore evaluate network topology, switching capacity, latency, congestion management and interconnect compatibility at the same time they select accelerators.

Power and Cooling Are Now Procurement Questions

The transition to rack-scale AI systems also changes the physical requirements of the data center.

Traditional enterprise servers could often be added gradually to existing facilities. Modern AI racks can require substantially more power and place much greater demands on cooling infrastructure.

The Rubin and Rubin Ultra roadmap makes facility planning especially important.

Future deployments must consider whether the data center can provide sufficient electrical capacity, coolant distribution, rack-level cooling, redundancy and serviceability.

This means a company considering a next-generation NVIDIA deployment should involve facilities and infrastructure teams early in the purchasing process.

Waiting until the GPUs arrive to ask whether the data center can support them is increasingly unrealistic.

Should Buyers Purchase Blackwell Now or Wait for Rubin?

For many organizations, this may be the most important question created by the AWS announcement.

There is no universal answer.

Blackwell can make sense when deployment is required now

Organizations with production workloads, available budgets and an immediate requirement for accelerated compute may gain more value from deploying current-generation infrastructure than delaying a project solely to wait for a future architecture.

Available Blackwell-generation systems also benefit from a more established server, networking and software ecosystem.

Time-to-compute matters.

A GPU that can deliver useful production work this quarter may create more business value than a theoretically superior system that will not fit an organization’s deployment timeline.

Rubin becomes compelling for new AI-factory designs

Organizations designing entirely new infrastructure have a different decision.

If the deployment timeline aligns with Vera Rubin availability, designing the facility around next-generation CPU, networking, storage, cooling and rack architecture may reduce the need for expensive infrastructure changes later.

Rubin is particularly relevant for buyers planning very large training environments, high-volume inference, agentic AI, scientific computing and other workloads where total AI-factory efficiency matters more than the performance of one GPU.

Rubin Ultra is a longer-term infrastructure decision

Rubin Ultra pushes this concept further.

NVIDIA has described future Rubin Ultra architectures capable of connecting hundreds of GPUs within extremely large NVLink domains.

At that scale, procurement becomes a data-center engineering project.

Electrical distribution, liquid cooling, networking, rack architecture, storage and facility design may need to be planned years before the hardware reaches production deployment.

For most ordinary enterprise buyers, waiting specifically for Rubin Ultra would therefore make little sense unless the organization is already planning infrastructure at hyperscale.

AWS Is Also Expanding Blackwell Capacity

The announcement does not mean AWS is abandoning Blackwell while waiting for Rubin.

AWS and NVIDIA also announced expanded Blackwell capacity, including infrastructure using the NVIDIA RTX PRO 4500 Blackwell Server Edition GPU for Amazon EC2 G7 instances.

The companies say G7 instances provide up to 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 generation. These figures are vendor-reported comparisons and actual application performance will depend on workload and configuration.

This simultaneous investment in Blackwell and Rubin shows why buyers should avoid thinking about GPU generations as a simple replacement cycle.

Different architectures can coexist because different workloads have different requirements.

Vera CPUs Are Another Important Part of the Announcement

GPUs receive most of the attention in AI infrastructure, but AWS and NVIDIA are also working to bring NVIDIA Vera CPU-based infrastructure to AWS.

This is particularly relevant to agentic AI.

An AI agent may need to execute code, call external tools, search databases, manage memory, process data and orchestrate several services before a GPU performs the next model operation.

Those workloads place significant demands on CPUs.

The result is another change in AI-system design: organizations need to calculate GPU and CPU requirements together rather than treating host processors as secondary components.

What This Means for Enterprise GPU Buyers

The AWS-NVIDIA expansion provides several useful lessons for organizations purchasing their own hardware.

First, do not choose an accelerator solely from benchmark results. Evaluate the workload, memory requirements, deployment schedule, software stack and infrastructure around it.

Second, budget for networking from the beginning. In larger GPU clusters, networking is part of compute performance.

Third, evaluate facility readiness before ordering high-density systems. Power and cooling limitations can become harder constraints than GPU availability.

Fourth, calculate total cost per useful workload, not simply acquisition price. Utilization, electricity, cooling, administration, networking and model throughput all influence real economics.

Finally, design for an upgrade path. The rapid transition from Blackwell to Rubin and eventually Rubin Ultra means infrastructure purchased today should ideally accommodate future generations without requiring a complete facility redesign.

NVIDIA Hardware Available for Buyers Today

Organizations that do not operate at AWS scale still have access to a wide spectrum of NVIDIA infrastructure.

Current enterprise options range from professional Blackwell GPUs and individual accelerators to DGX systems, high-speed networking components and rack-scale platforms.

For buyers comparing infrastructure, it can be useful to examine categories separately:

Professional and server GPUs for workstation, visualization and specialized compute deployments.

Data-center accelerators and DGX systems for training, inference and high-performance computing.

SuperNICs, DPUs and switches for multi-GPU and multi-node networking.

Rack-scale systems for organizations building dedicated AI factories.

AI Robot Supplier’s catalog currently includes NVIDIA infrastructure across these categories, including Blackwell-generation professional GPUs, DGX systems, ConnectX networking, BlueField technology and emerging Vera Rubin platforms.

The appropriate configuration should be determined from the intended workload rather than simply selecting the most expensive or newest product.

The Bigger Picture: AI Infrastructure Is Becoming Full Stack

The significance of AWS’s planned 2-million-GPU expansion is not simply the number of accelerators involved.

It demonstrates how quickly AI infrastructure is becoming vertically integrated.

The accelerator, CPU, memory, network, storage, cooling system, rack architecture and software environment increasingly need to be engineered together.

For hyperscalers such as AWS, that means building entire AI factories.

For enterprise buyers, the same principle applies at a smaller scale.

A four-GPU workstation, an eight-GPU server and a 100-rack cluster are very different projects, but each performs best when the surrounding infrastructure is designed for the accelerator rather than added afterward.

What Buyers Should Watch Next

The most important developments to watch during the next 12 to 24 months will be actual Rubin system availability, cloud pricing, Rubin Ultra deployment timelines, networking architectures, power requirements and real-world workload economics.

Independent production benchmarks will also become increasingly important.

Manufacturer performance claims are useful for understanding the direction of the technology, but purchasing decisions should ultimately be based on the buyer’s own workload, deployment environment and total cost of ownership.

The AWS announcement confirms that demand for accelerated AI infrastructure remains exceptionally strong.

It also confirms something more important for procurement teams: the AI infrastructure buying decision is no longer just about which GPU to purchase. It is about designing the complete system around it.

Frequently Asked Questions

How many additional NVIDIA GPUs does AWS plan to deploy?

AWS and NVIDIA announced plans to deploy 2 million additional NVIDIA GPUs during 2027 and 2028 across AWS’s global infrastructure.

Which NVIDIA GPU generations are included?

The companies specifically named NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs.

Is NVIDIA Rubin available now?

NVIDIA says the Vera Rubin platform has entered production, but availability of specific systems, cloud instances and configurations varies by partner and deployment schedule. Buyers should confirm actual delivery timelines before basing infrastructure plans on a particular system.

What is the difference between Rubin and Blackwell?

Blackwell is NVIDIA’s current-generation accelerated-computing architecture, while Rubin is the next major platform generation. Vera Rubin combines the Rubin GPU with Vera CPUs, NVLink 6, ConnectX-9, BlueField-4 and next-generation networking technologies.

Should I wait for Rubin instead of buying Blackwell?

Not automatically. Organizations requiring compute today may achieve greater value by deploying available Blackwell infrastructure. Buyers constructing new large-scale AI facilities with later deployment dates may have stronger reasons to design around Vera Rubin.

Why does networking matter for AI GPU clusters?

Large AI workloads distribute computation across multiple accelerators. If data cannot move between those accelerators quickly enough, network congestion and latency can reduce GPU utilization and extend job completion times.

Where can businesses compare NVIDIA GPU and AI infrastructure options?

Businesses should compare the complete configuration—including accelerator, memory, server architecture, networking, power, cooling, software and expected delivery timeframe—rather than evaluating GPU specifications alone. AI Robot Supplier provides NVIDIA hardware and AI infrastructure categories that can be compared according to deployment requirements.

Related reading

Leave a Reply

Your email address will not be published. Required fields are marked *