GPU Cloud Pricing Comparison 2026: H100, A100, RTX 4090

GPU Cloud Pricing Comparison 2026: H100, A100, RTX 4090 (Updated)

This GPU cloud pricing comparison is updated quarterly to reflect current market rates for renting GPU compute. Prices change frequently as new providers enter the market and hardware availability shifts. The figures below were last verified in April 2026.

GPU Cloud Pricing Comparison: H100 Instances

ProviderGPUVRAMPrice/hr (on-demand)Notes
RunPod Secure CloudH100 SXM580 GB~$2.49Datacenter, high uptime
Lambda LabsH100 SXM580 GB~$2.49–3.50Cluster available
CoreWeaveH100 SXM580 GB~$2.65–3.00Enterprise, NVLink
Vast.aiH100 SXM580 GB~$1.89–2.20Host-dependent reliability
Google Cloud (on-demand)H100 A380 GB~$5.50+Spot much cheaper
AWS P4d instancesA100 (via p4d)40 GB×8~$32/hr (8-GPU)$4/GPU/hr equiv.

GPU Cloud Pricing Comparison: A100 Instances

ProviderGPUVRAMPrice/hrNotes
RunPodA100 SXM480 GB~$1.64Managed, reliable
Lambda LabsA100 SXM480 GB~$1.99Premium datacenter
VultrA100 PCIe80 GB~$2.40US/EU regions available
Vast.aiA100 (mixed)40/80 GB~$1.20–1.50Variable quality
DigitalOceanH100 (new)80 GB~$3.99+Simplified management

GPU Cloud Pricing Comparison: RTX 4090 Instances

ProviderGPUVRAMPrice/hrNotes
RunPod Community CloudRTX 409024 GB~$0.74Consumer GPU, good uptime
Vast.aiRTX 409024 GB~$0.30–0.45Cheapest available, variable
HetznerRTX 4000 Ada20 GB~$0.80EU datacenter, stable

Best Price per VRAM: Which GPU Gives the Most Memory per Dollar?

For workloads that need raw VRAM (large model inference, multi-LoRA serving), the A100 80 GB at ~$1.64/hr on RunPod gives 80 GB for $1.64 — or $0.0205/GB/hr. The RTX 4090 at $0.74/hr gives 24 GB for $0.0308/GB/hr. The A100 wins on cost-per-GB for VRAM-heavy workloads.

H200 and B200 Pricing: The New Tier for 2026

The GPU cloud pricing landscape expanded significantly in 2026 with wider availability of H200 and Blackwell B200 instances. The H200 (141 GB HBM3e) is available on RunPod Secure Cloud at $3.59/hr and Lambda Labs at $3.29/hr. The B200 (192 GB HBM3e, FP4 support) appears on RunPod at $5.98/hr and Nebius at $5.50/hr. These higher-tier GPUs cost more per hour but deliver significantly more tokens per second for 70B+ inference workloads, which can result in lower effective cost per million tokens — particularly for teams running sustained high-volume endpoints. For a full cost-per-token analysis across H100, H200, and B200, see our H200 vs B200 vs H100 cost per token comparison.

Spot and Reserved Pricing: Reducing the Bill by 50–90%

Spot GPU instances represent the largest available discount in cloud compute. Spheron lists H100 SXM spot at $1.03/hr (versus $2.49 on-demand — a 59% discount) and B200 spot at $2.12/hr (versus $4.99 on-demand). GCP A3 spot instances list at $2.25/hr. Spot instances are preemptible with short notice, making them unsuitable for serving live inference endpoints but highly cost-effective for training runs with checkpoint support. Budget roughly 10% of your spot spend for jobs that require restart due to preemption — the math still strongly favors spot for training at any meaningful scale.

Reserved GPU contracts (1-year or 3-year) offer a middle path. H100 reserved pricing has dropped to approximately $1.70/hr across specialist providers. B200 reserved is available at $2.25/hr. Reserved contracts work best for teams with predictable baseline compute demand and at least 6 months of data to back the utilization forecast.

For a complete breakdown of which provider to choose based on your workload, see our best GPU VPS for AI guide. For head-to-head provider comparisons, see RunPod vs Vast.ai vs Lambda Labs.

How We Collected These Prices

All prices are sourced directly from provider pricing pages and verified via test account creation. Spot/preemptible prices are not included in the main tables — on-demand only. Prices are in USD and exclude storage costs. Last verified: June 2026.

Sources

Related reading: owned clusters vs cloud

When sustained utilization exceeds ~20%, on-premise economics often beat reserved cloud. See GPU cluster TCO 2026 and power infrastructure for US clusters. For an alternative to renting GPUs entirely, see custom AI silicon vs GPU rental 2026, covering TPU, Trainium, and Maia economics.

Iovanny Olguín Ávila
Author: Iovanny Olguín Ávila

Computer Systems Engineer with an MSc in Computer Science. I apply quantitative analysis and data-driven methodologies to evaluate financial instruments, investment vehicles, and emerging technologies. My technical background allows me to cut through marketing language and analyze the actual mechanics of financial products — from HELOC structures to Medicare Advantage plan design to business credit card reward algorithms.

Leave a Comment