best cloud GPU for AI

Best Cloud GPU for AI: 8 Providers Compared on Price (2026)

READING TIME 13 min read |
LAST UPDATED September 21, 2026 |
CATEGORY Comparisons |
ARTICLE VIEWS 2 views

Best cloud GPU for AI work still starts with one number: what an hour of compute costs you. Renting an NVIDIA H100 averaged about $3.84 per GPU-hour across 27 tracked providers in mid-September 2026 — but the cheapest advertised rates sit near $1.47 on marketplaces and the most expensive hyperscaler listings reach $10 or more. Same chip, wildly different bill. This guide compares the eight providers that matter right now, with pricing snapshots from September 2026, so you can match the right cloud to your workload instead of paying for hardware you don’t need.

How cloud GPU rental actually works

Strip away the marketing and there are really only three pricing models in this market, plus three kinds of seller.

On-demand is the default: a published hourly rate, the instance yours until you stop it. Spot (also called preemptible) instances run on spare capacity at deep discounts — commonly 60–77% off — but the provider can reclaim the machine with little warning. Fine for checkpointed training, fatal for anything that can’t be interrupted. Reserved pricing trades a commitment, anywhere from a month to three years, for 30–60% off on-demand.

The sellers fall into three camps. Hyperscalers (AWS, Google Cloud, Azure) sell GPUs inside their giant cloud ecosystems — expensive per hour, but with compliance certifications, private networking, and the tooling enterprises already run on. Neoclouds (Lambda Labs, CoreWeave, RunPod, Nebius, Crusoe Cloud) are GPU-first companies built around AI workloads; they typically undercut hyperscalers by 30–60% on the same NVIDIA hardware. Marketplaces (Vast.ai, Spheron) let independent hosts list spare GPUs — the cheapest option by far, but with variable reliability and no meaningful support.

Prices are quoted per GPU-hour and move constantly. One independent index tracking 27 providers put the volume-weighted H100 average at $3.84 an hour on September 13, 2026 — up about 6% over the previous 90 days. Treat every figure below as a September 2026 snapshot, and check live pricing before committing to a long run.

The 8 providers, head to head

All prices below are on-demand per-GPU-hour, drawn from published provider comparisons in September 2026. Marketplace figures are ranges because hosts set their own prices. The H200 column is new: NVIDIA’s 141 GB card has become the default recommendation for memory-bound inference this year.

ProviderTypeH100 80GB ($/hr)H200 141GB ($/hr)A100 80GB ($/hr)Min billingMulti-GPUBest for
Lambda LabsNeocloud$3.29–$4.29$1.29–$2.791 hourYes (8x)Simple, ML-ready single-GPU and small-cluster work
RunPodNeocloud$2.39–$2.99~$4.59$1.09–$1.791 minuteYes (8x)Iterative experiments and budget-conscious training
Vast.aiMarketplace$1.40–$2.10~$2.25+$0.70–$1.501 minuteLimitedCheapest possible prototyping
CoreWeaveNeocloud$2.23–$6.16~$6.31$1.02–$2.7010 minutesYes (256+)Enterprise-scale distributed training
NebiusNeocloud~$3.851 minuteYesEuropean data-residency workloads
Crusoe CloudNeocloud~$3.90YesAMD-based workloads; clean-energy positioning
AWSHyperscaler$3.22–$6.88~$7.91~$3.431 secondYes (EFA)Teams already living on AWS
Google CloudHyperscaler$3.06–$11.68~$10.60$1.54–$2.481 minuteYesTPU workloads; GCP-native teams

Two patterns jump out immediately. First, the spread on the same card is enormous — an H100 can cost $1.40 an hour or nearly $12 depending on where you rent it. Second, the hyperscalers charge roughly double the neoclouds for NVIDIA’s own silicon, which is why the neoclouds have kept eating their lunch since the GPU shortage made everyone price-sensitive.

Provider by provider: who each one is for

Lambda Labs: the straightforward default

Lambda built its name selling GPU workstations, and its cloud reflects that: transparent on-demand pricing, machines arriving with PyTorch and the usual ML stack preinstalled, and no egress fees. Listed H100 rates ran $3.29–$4.29 an hour in published September 2026 comparisons, with A100s around $2.79 for 8-GPU nodes. The tradeoff is availability — a handful of US data centers — and no exotic networking for giant training runs. For researchers and small teams who want to train in minutes without learning Kubernetes, it’s the least painful on-ramp.

CoreWeave: the enterprise neocloud

CoreWeave is the heavyweight of the neocloud world: GPU-dense fleets, InfiniBand networking, Kubernetes-native orchestration, and contracts built for companies training at real scale. Its listed H100 rate is $6.16 an hour on the enterprise on-demand tier (reserve-tier pricing runs lower), and it was among the first clouds to offer NVIDIA’s Blackwell cards at scale — B200s listed around $8.60 an hour and H200s around $6.31 in September 2026. Choose CoreWeave when you need hundreds of GPUs networked together with enterprise SLAs — not when you need one card for a weekend.

RunPod: the experimenter’s workbench

RunPod’s killer feature is per-second billing with one-minute minimums — ideal for the loop most ML developers live in: spin up, run an experiment, tear down, repeat. Its H100 listed at $2.89 an hour on its vetted Secure Cloud tier in early September 2026, with spot cheaper still, and an RTX 4090 starts near $0.34 an hour. It also lists AMD’s MI300X at roughly $1.89 an hour and NVIDIA’s H200 near $4.59, a sign of where the market is heading as memory-hungry inference work grows. RunPod is the best value for iterative fine-tuning and inference prototyping.

Vast.ai: the bargain basement

Vast.ai is a marketplace where independent hosts rent out spare GPUs, from gaming PCs to data-center racks — which produces the lowest prices anywhere: H100s at $1.40–$2.10 an hour, spot listings from about $1.47, A100s under a dollar, RTX 4090s from $0.35. The catch is reliability: hosts can disappear mid-job, there’s no support line, and multi-GPU setups are limited. Checkpoint aggressively and treat it as disposable compute. For students, hobbyists, and anyone whose training run can survive an interruption, nothing else comes close on price.

Nebius and Crusoe Cloud: the new breed

Nebius (the former Yandex cloud business, rebuilt for AI infrastructure) prices its H100s around $3.85 an hour and is one of the few serious options with a European footprint — relevant if data residency matters. Crusoe Cloud is the newest name here, notable for its modular AI data centers and for listing AMD’s MI300X at about $1.71 an hour, the cheapest flagship-accelerator rate in any published comparison. Both are worth watching if you want NVIDIA alternatives or compliance coverage the US neoclouds don’t offer. For background on the physical infrastructure behind these providers, see our explainer on what AI data centers are.

AWS and Google Cloud: the enterprise tax

Renting an H100 on AWS runs $3.22–$6.88 an hour per GPU; Google Cloud’s A3 Mega instances list at about $11.68 per GPU-hour on-demand (though committed-use discounts cut that sharply). You’re paying for SOC 2 and HIPAA compliance, private networking, identity management, and petabyte-scale storage integration — not for the GPU itself. If your company already runs on a hyperscaler and moving data out would be its own project, the premium can be rational. For everyone else, it’s roughly double the neocloud price for the same silicon.

Which GPU should you actually rent?

The provider matters less than the card. Here’s how the current lineup maps to real jobs, with September 2026 pricing snapshots and the rough VRAM math that decides whether your model fits:

GPUVRAMTypical cloud rateWhat it’s for
RTX 409024 GB$0.34–$0.59/hrInference and QLoRA fine-tuning up to ~30B parameters; the budget sweet spot
L40S48 GB$0.92–$1.19/hrThe value pick: most of an A100’s practical usefulness at a fraction of the cost
A100 80GB80 GB$1.02–$2.79/hrFine-tuning up to 70B models; the most cost-efficient data-center card
H100 80GB80 GB$1.40–$6.88/hrSerious training and large-scale inference; 2–2.5x faster than A100 on modern workloads
H200 141GB141 GB$2.25–$10.60/hrMemory-bound inference: one H200 replaces two H100s for a 70B model in FP16; same 700W power as the H100
B200 180GB180 GB$3.74–$8.60/hrCutting-edge training and FP4 inference; long-term reservations drop as low as ~$2.25/hr
MI300X (AMD)192 GB$1.71–$1.89/hrMemory-bound inference; competitive for ROCm-compatible workloads at a lower price

The H200 deserves the extra attention it’s getting this year. With 141 GB of HBM3e memory — nearly double the H100’s 80 GB at the same 700W power draw — it’s a drop-in rack replacement that triples the effective KV-cache headroom for serving large models. Independent trackers put it at a 15–40% premium over the H100 on neoclouds ($4.43 average vs $3.84 in September 2026), but for inference that premium often pays for itself: a single H200 can serve a 70B model in FP16 where two H100s were needed, cutting inter-GPU overhead and simplifying the whole serving setup.

With QLoRA — the technique that quantizes a model to 4-bit precision and trains small adapter layers instead of the whole model — a 7–8 billion parameter model needs only 8–12 GB of VRAM, so even a modest card works. A 70B model at 4-bit fits in roughly 38 GB, which is why the A100 80GB has become the standard recommendation for open-model fine-tuning: one published analysis puts a full 70B QLoRA run at $34–$51 and 24–36 hours on an A100. Full fine-tuning (updating every weight, no adapters) roughly doubles or triples the memory requirement — a 70B model at full precision wants around 420 GB of VRAM, which means multiple H100s or A100s and a cloud bill to match.

Data-center GPUs are a different species from gaming cards. An RTX 4090 tops out at 24 GB of GDDR memory; an H100 carries 80 GB of much faster HBM memory with error correction, NVLink for multi-GPU scaling, and hardware support for the FP8 precision format modern training relies on. Consumer cards are fine for learning and small jobs. Production training wants data-center silicon.

Rent or buy: the breakeven math

An H100 costs roughly $25,000–$40,000 to buy new, and the instinct to just own the thing is understandable. But every independent total-cost-of-ownership study lands in the same place: ownership only wins when the GPU is actually working most of the time.

The crossover sits between 40% and 77% sustained utilization, depending on hardware and how honestly you count power, cooling, networking, and 15–20% annual depreciation. Below that line, renting wins; above it, buying or a long-term reserved commitment wins. Most organizations that measure real utilization for the first time land at 35–55% — squarely in rent territory.

For consumer cards the math is friendlier to ownership. One three-year comparison found that running an RTX 4090 eight hours a day costs about $3,456 in cloud rental versus roughly $2,073 to own including electricity — a saving of nearly $1,400. But that assumes eight hours a day, every day, for three years. For bursty workloads — a training run here, an experiment there — renting a card for one run and shutting it down remains the cheaper move.

What this means practically

Starting out, you need far less than you think. A single RTX 4090-class card — rented for under $0.50 an hour — handles most fine-tuning and inference experiments on open models up to 30B parameters. Scale up to an A100 when you’re working with 70B models or need the memory headroom; reach for the H100 only when training speed is the bottleneck, and consider the H200 when memory — not raw compute — is what’s slowing your serving down.

On providers: start with Lambda or RunPod for simplicity, drop to Vast.ai when a run is cheap enough that an interruption wouldn’t hurt, and graduate to CoreWeave or a hyperscaler when you need scale, networking, or compliance. And don’t sign a year-long commitment before measuring how many GPU-hours you actually burn in a month — this market rewards the uncommitted.

If you’re renting GPUs to power AI agents or research workloads, our guides to AI agents for scientific research and AI coding tools for beginners cover what to actually do with the compute once you’ve rented it.

FAQ

Which cloud GPU provider is the cheapest?

Vast.ai is consistently the cheapest in published comparisons — H100s at roughly $1.40–$2.10 an hour, with spot listings dipping to about $1.47 in September 2026 — because it’s a marketplace of independent hosts rather than a company with its own data centers. RunPod is usually the cheapest option with vetted, reliable hosts. The tradeoff is reliability: marketplace hosts can vanish mid-job, so checkpoint your work.

How much does it cost to rent an H100 per hour?

As of September 13, 2026, an independent index tracking 27 providers put the volume-weighted H100 average at $3.84 per GPU-hour. Neoclouds list from about $2.89 (RunPod) to $3.29 (Lambda), CoreWeave lists $6.16 on its enterprise tier, AWS runs $3.22–$6.88, and Google Cloud lists up to about $11.68 per GPU-hour on its A3 Mega instances. Spot and marketplace rates go lower; hyperscaler rates go higher. Prices move frequently, so check live rates before a long run.

How much does it cost to rent an H200 per hour?

About $4.43 per GPU-hour on average across 20 tracked providers in September 2026, roughly a 15–40% premium over the H100 on neoclouds. Listed rates range from around $2.25–$2.45 for 8-GPU nodes at smaller specialists to $4.59 at RunPod, $6.31 at CoreWeave, $7.91 on AWS, and about $10.60 on Google Cloud’s A3 Ultra instances. For memory-bound inference, the premium often pays for itself since one H200 replaces two H100s.

Should I rent an H100 or an H200?

Pick the H200 if your workload is memory-bound — large-model inference, long-context serving, big KV caches. Its 141 GB of HBM3e (vs 80 GB on the H100, at the same 700W power draw) triples effective KV-cache headroom, and a single H200 can serve a 70B model in FP16 where two H100s were needed. Pick the H100 for raw training throughput and wider availability: it’s listed at 36+ providers, is cheaper per hour, and remains the reference card for large training runs.

How much does it cost to rent a B200 per hour?

About $6.39 per GPU-hour on average across 12 tracked providers in September 2026. Specialist listings range from $3.74 for 8-GPU nodes at smaller providers to $6.79 at RunPod, $6.99 at Lambda, and $8.60 at CoreWeave, with Oracle listing as high as $14. Long-term reserved contracts (36 months) have dropped as low as roughly $2.25 an hour. The B200 delivers about 2.5x the H200’s inference throughput and supports FP4 precision, but needs liquid cooling.

Is it cheaper to rent or buy a GPU for AI work?

It depends on utilization. Independent cost studies put the breakeven between 40% and 77% sustained GPU usage: below that, renting wins; above it, buying or a long-term reserved commitment wins. Most teams that measure real utilization land at 35–55%, in renting territory. For consumer cards like the RTX 4090 used several hours daily, buying can save over $1,000 across three years.

What GPU do I need to fine-tune an LLM?

With QLoRA (4-bit quantization plus adapter training), a 7–8B model needs only 8–12 GB of VRAM — a 12GB card works, and a 24GB RTX 4090 is comfortable. A 70B model at 4-bit needs about 38 GB, making the A100 80GB the standard pick (a full 70B QLoRA run costs roughly $34–$51 in cloud rental). Full fine-tuning without adapters roughly doubles or triples these requirements.

Lambda Labs vs RunPod: which is better?

They serve different users. Lambda Labs is simpler — transparent hourly pricing ($3.29 for an H100 in published September 2026 lists), ML-ready machines, no egress fees — ideal for individuals and small teams doing single-GPU and small-cluster work. RunPod is cheaper ($2.89 for an H100) with per-second billing and one-minute minimums, built for the spin-up, experiment, tear-down loop. Pick Lambda for getting started fast; pick RunPod for iterative experimentation.

Can I use a consumer GPU like the RTX 4090 for AI training?

Yes, for many jobs. The 24GB RTX 4090 handles inference and QLoRA fine-tuning on models up to about 30B parameters, and rents for as little as $0.34 an hour. Its limits are real, though: 24GB caps full-precision fine-tuning of larger models, and consumer GDDR memory is far slower than the HBM memory on data-center cards. For production training of large models, data-center GPUs remain the practical choice.

Are cloud GPU prices going up or down?

Slightly up, lately. An independent GPU price index showed the H100 average rising from $3.62 to $3.84 per GPU-hour over the 90 days to mid-September 2026 — a 6.1% increase — though the 30-day move was a small decline of 0.4%. H200 prices fell 2.4% over the same recent week. The shortage-era spikes are over, but don’t plan a budget on the assumption that prices only fall.

References

  1. Mercatus — “GPU Rental Prices: H100, H200, B200 Cost Per Hour by Provider” — mercatus-ai.com — September 2026
  2. Spheron — “GPU Rental Price Comparison: RunPod, Lambda, CoreWeave (read 8 Sep 2026)” — spheron.network — September 2026
  3. GPUaaS — “H100 vs H200 vs B200: Which GPU to Rent in 2026” — gpuaas.com — 2026
  4. GPUSmith — “GPU Cloud Rental Prices 2026: H100 vs H200 Cost Comparison” — gpusmith.com — 2026
  5. Lyceum Technology — “Lambda Labs vs RunPod vs Vast.ai: GPU Cloud Comparison” — lyceum.technology — August 2026
  6. GPU Smith — “Own vs Rent GPUs: On-Prem vs Cloud AI Cost Comparison 2026” — gpusmith.com — July 2026
  7. Thunder Compute — “Best GPU for AI in 2026: Local Hardware and Cloud Options” — thundercompute.com — September 2026
  8. gpu.fm — “Cloud GPU Providers Compared (2026): Lambda, CoreWeave, RunPod & More” — gpu.fm — 2026

Comments

One response to “Best Cloud GPU for AI: 8 Providers Compared on Price (2026)”

  1. […] internal use — it favors cloud for sporadic or experimental work. Our guide to the best cloud GPUs for AI covers the other side of that […]

Leave a Reply

Your email address will not be published. Required fields are marked *