Comparison

RTX 4090 vs H100: Cost per FLOP for AI training.

Live rental prices and spec-sheet tensor throughput give the real cost per petaflop-hour on each card — and why memory and interconnect change the answer.

By Vi Nguyen · Published · Updated · Prices on this page are live · Editorial policy

RTX 4090 — cheapest right now
$0.30/hr
450 W · 24 GB GDDR6X · ~165 TFLOPS BF16 dense
See all providers →
H100 — cheapest right now
$2.29/hr
Up to 700 W (SXM) · 80 GB HBM3 · ~989 TFLOPS BF16 dense (SXM)
See all providers →

The RTX 4090 and the H100 sit at opposite ends of the GPU rental market. One is a gaming card that became the default budget choice for AI work; the other is the datacenter accelerator most large training runs of the last few years were built on. On a price-per-FLOP basis the gap between them is much smaller than the gap in hourly price suggests, and sometimes it runs the other way. This article works out the real cost per teraflop-hour on each card from live rental prices, then explains why FLOPS alone is the wrong basis for the decision, and what to use instead.

The headline number lies

Spec sheets make this comparison look lopsided. NVIDIA lists the H100 SXM at about 989 TFLOPS of dense BF16 tensor throughput, and nearly 2,000 with sparsity. The RTX 4090's most-quoted number is about 83 TFLOPS, but that is its shader (non-tensor) FP16/FP32 rate, which isn't what matrix-heavy AI code uses. Compare the wrong pair of numbers and you'd conclude the H100 is twelve to twenty-four times faster. So why does anyone train on 4090s?

Because the price gap is also large, and on the right metric it's in the same range. The 4090 launched at a $1,599 MSRP and is widely available on host-supplied marketplaces such as Vast.ai and Clore.ai. The H100 is a datacenter part costing tens of thousands of dollars, and it rents for several times more per hour even at the cheapest providers. Right now the cheapest 4090 in our data is $0.30/hr on io.net, and the cheapest H100 is $2.29/hr on Theta EdgeCloud. Divide one by the other and you get the number that actually matters: how much compute each dollar buys.

Dense vs sparse TFLOPS, and which number to use

Before you can compare cost per FLOP, you need to compare the same kind of FLOP. GPU spec sheets list several numbers for each card, and they differ by 2× or 4× depending on which footnote you read.

  • Shader vs tensor throughput. The general-purpose CUDA cores do ordinary floating-point math. Tensor cores do matrix multiplies, which is what transformer training and inference spend almost all their time on. The 4090's 83 TFLOPS figure is shader throughput. For AI work, compare tensor throughput.
  • Dense vs sparse. Since the A100, NVIDIA has quoted a "with sparsity" figure that is exactly twice the dense one. It assumes 2:4 structured sparsity: in every group of four weights, two are zero, and the hardware skips them. Almost no training run and few inference deployments use this, because the model has to be pruned to that pattern and quality usually suffers. For honest comparisons, use the dense number. Sparsity isn't a feature one of these cards has and the other lacks: both Ada (the 4090) and Hopper (the H100) support it.
  • Accumulation precision. On GeForce cards, NVIDIA halves tensor throughput when the multiply accumulates in FP32, which is what standard mixed-precision training uses. The 4090's dense FP16/BF16 tensor rate is about 330 TFLOPS with FP16 accumulation and about 165 TFLOPS with FP32 accumulation. Datacenter cards don't have this restriction. For training, 165 is the fair number.
  • FP8. Both cards have FP8 tensor cores (fourth-generation tensor cores on Ada and Hopper), and both roughly double their peak throughput over BF16 at FP8. The H100 SXM is rated at about 1,979 dense FP8 TFLOPS. The software ecosystem for FP8 training on the H100, including NVIDIA's Transformer Engine, is more mature, but the 4090 is not emulating FP8.

So the like-for-like comparison for mixed-precision training is about 165 dense BF16 TFLOPS for the RTX 4090 against about 989 for the H100 SXM, a ratio of roughly 6×. The PCIe version of the H100 runs at lower clocks and power, so its rated throughput is lower. Our generic H100 listing groups offers that don't specify the form factor; if you need SXM specifically, check the H100 SXM and H100 PCIe pages.

Doing the actual math

Cost per teraflop-hour is the hourly price divided by the dense BF16 tensor throughput. Because the numbers get small, it's easier to read per petaflop-hour (1,000 TFLOP-hours of sustained compute).

  1. RTX 4090: $0.30/hr ÷ 165 TFLOPS × 1,000 = $1.82 per petaflop-hour.
  2. H100 (SXM rating, dense): $2.29/hr ÷ 989 TFLOPS × 1,000 = $2.32 per petaflop-hour.
  3. Ratio: at today's cheapest listings, the RTX 4090 is about 1.3× cheaper per peak FLOP.

The break-even point is easy to remember: because the H100 SXM has about six times the dense BF16 throughput, it is cheaper per peak FLOP whenever its hourly price is less than about six times the 4090's (989 ÷ 165 ≈ 6.0). Cheap marketplace 4090 listings tend to keep the 4090 ahead on this measure; cheap H100 listings can close the gap entirely.

But peak FLOPS is not what you get. The next step is to adjust for how much of that peak a real job sustains. Suppose, hypothetically, a single-GPU fine-tuning job sustains 35% of peak on either card. The per-FLOP ratio doesn't change. Now suppose you spread a job across eight 4090s and communication overhead drops their effective utilization to 20%, while eight NVLink-connected H100s hold 35%. The 4090's effective cost per useful FLOP rises by 35 ÷ 20 = 1.75×, which can be enough to flip the result. Raw FLOPS isn't what you're paying for. You're paying for the ability to actually finish a training run.

Memory capacity and interconnect matter as much as FLOPS

FLOPS decide how fast a job runs once it fits. Memory decides whether it fits at all, and bandwidth decides how much of the FLOPS you can use.

Capacity. The 4090 has 24 GB; the H100 has 80 GB. Using the site's rule of thumb (parameters × bytes per parameter, plus KV cache and 10–20% overhead), a 24 GB card holds a 7B–8B model at FP16 for inference, around 13B at 8-bit, and roughly 30B at 4-bit. A 70B model at 4-bit needs about 35 GB for weights alone, so it doesn't fit on one 4090. On an 80 GB H100, a 70B model fits at 8-bit (about 70 GB, tight once you add KV cache) or comfortably at 4-bit; at FP16 it needs about 140 GB, which means two H100s. For training the gap is bigger, because a full fine-tune with Adam needs about 16 bytes per parameter. The VRAM fit calculator runs these numbers for specific models.

Memory bandwidth. The 4090's GDDR6X delivers about 1 TB/s. The H100 SXM's HBM3 delivers about 3.35 TB/s, and the PCIe version about 2 TB/s. For LLM inference, where each generated token requires reading all the active weights, decode speed is roughly bounded by bandwidth ÷ bytes of weights. An 8 GB model (8B parameters at 8-bit) has a theoretical ceiling of about 1,000 ÷ 8 ≈ 125 tokens per second per sequence on a 4090, and about 3,350 ÷ 8 ≈ 420 on an H100 SXM. These are upper bounds, not measurements, but the ratio between them (about 3.3×) is smaller than the FLOPS ratio (about 6×). For single-stream inference, that makes the 4090 look better than a FLOPS comparison would suggest.

Interconnect. This is where the cards differ most. An H100 SXM has NVLink 4 with 900 GB/s of total GPU-to-GPU bandwidth, and in an 8-GPU HGX server every GPU reaches every other at that speed through NVSwitch. The 4090 has no NVLink connector. Multi-4090 systems communicate over PCIe, and a PCIe 4.0 x16 slot provides about 32 GB/s in each direction, more than an order of magnitude less. On top of that, according to widespread reports, NVIDIA does not enable peer-to-peer transfers between GeForce 40-series cards in its standard driver, so traffic between two 4090s generally goes through host memory, which adds latency and overhead.

What the 4090 can't do

  • VRAM ceiling: 24 GB caps a single 4090 at about 7B–8B models at FP16, 13B at 8-bit, or around 30B at 4-bit for inference. The H100's 80 GB holds a 70B model at 8-bit or 4-bit on one card.
  • No NVLink: without it, 8× 4090s can't behave like one large GPU. Tensor parallelism, which splits each layer across GPUs and communicates on every layer, is effectively off the table at useful speeds.
  • Memory bandwidth: about 1 TB/s against the H100's 2–3.35 TB/s, which limits large-batch inference and any workload that streams weights.
  • Datacenter features: no MIG partitioning to split one card between several jobs, and consumer cards are generally hosted with less redundancy than datacenter GPUs. On marketplaces, reliability depends on the individual host.

Multi-4090 scaling: where it breaks down

Renting eight cheap cards to match one expensive node is the obvious idea, and for some workloads it works. The question is how much of the work requires GPUs to talk to each other.

Independent work scales perfectly. Image generation, batch inference of models that fit on one card, hyperparameter sweeps, and evaluation runs need no communication. Eight 4090s do eight times the work of one. This is where multi-4090 setups make the most sense.

Data-parallel training scales partially. Each GPU holds a full copy of the model and processes different data, and after every step the gradients are averaged across GPUs with an all-reduce. With a ring all-reduce, each GPU sends and receives about 2 × (N − 1) ÷ N times the gradient size per step. As an illustrative example, take a 1.5B-parameter model with 3 GB of BF16 gradients on 8 GPUs: each GPU moves about 2 × 7 ÷ 8 × 3 = 5.25 GB per step. At a hypothetical 20 GB/s of effective PCIe throughput, that's about 0.26 seconds per step spent communicating. At a hypothetical 400 GB/s of effective NVLink throughput, it's about 0.013 seconds. If the compute part of the step takes one second, the 4090 cluster spends a meaningful share of its time waiting, unless the framework overlaps communication with computation well. LoRA and QLoRA are the exception: their gradients are tiny, so data-parallel LoRA on several 4090s scales far better than full fine-tuning does.

Model-parallel training barely scales. Once a model doesn't fit on one card, you have to split it: tensor parallelism across each layer's matrices, pipeline parallelism across layers, or fully sharded data parallelism (FSDP or ZeRO-3), which gathers weights on every forward and backward pass. All of these multiply communication volume. Pipeline parallelism is the most PCIe-friendly, but it introduces idle "bubbles" while GPUs wait for each other. This is where NVLink earns its price, and it's the main reason large-model training happens on H100 nodes rather than 4090 racks.

There are practical limits too. Some marketplace hosts with several 4090s run cards on x8 or narrower PCIe links, and multi-GPU consumer boxes vary widely in motherboard, CPU and cooling. Check a host's reported PCIe bandwidth before you build a multi-GPU job around it.

Where the 4090 still wins

  • Image generation (Stable Diffusion XL, Flux): the models fit in 24 GB, each job runs on one card, and the 4090 has FP8 tensor cores for FP8 Flux. See the best GPUs for Stable Diffusion and Flux.
  • Hobbyist fine-tuning: LoRA and QLoRA on 7B–13B models are a natural 4090 job. The LoRA and QLoRA cost guide works through the memory math.
  • Batch inference of smaller models: serving a 7B or 8B model is well within a 4090's memory, and many independent replicas scale cleanly.
  • Single-host experimentation: when a job doesn't need to scale past one card, the 4090's per-FLOP advantage at typical marketplace prices is real, and short iterations matter more than peak speed.

When to pay for H100s

  • Full fine-tuning or training above a few billion parameters: at 16 bytes per parameter, even a 7B full fine-tune needs over 100 GB plus activations, so you need the memory and the multi-GPU bandwidth.
  • Serving 70B-class models on a single card at 8-bit or 4-bit, or at FP16 across two cards.
  • Production inference where you need consistent latency and a provider with a real SLA.
  • Fast multi-GPU scaling: tensor-parallel and fully sharded training depend on NVLink bandwidth that PCIe can't match.
  • FP8 training at scale: the H100's FP8 peak is double its BF16 peak, and the tooling around it is more mature than on consumer cards.

A decision framework

Run through these in order, and stop at the first one that settles it.

  1. Does the job fit on one 24 GB card? If not, and quantization isn't an option, you need more memory. Consider a 48 GB card (see our 48 GB comparison) or an 80 GB H100 before trying to split the job across 4090s.
  2. If it fits, is it independent work? Image generation, batch inference and sweeps: rent 4090s, as many as you need. Check the cost per petaflop-hour above; at typical marketplace prices, the 4090 usually wins.
  3. Is it multi-GPU training? If the gradients are small (LoRA, QLoRA), multi-4090 data parallelism is workable. If you need to shard the model, rent H100s with NVLink.
  4. Does wall-clock time matter? If a run that finishes in three hours instead of eighteen lets you iterate faster or ship sooner, the H100 can be worth paying for even when it costs more per FLOP.
  5. Do you need reliability guarantees? For production workloads, a managed provider such as Lambda or RunPod's Secure Cloud is usually worth a premium over a marketplace host. See GPU marketplaces vs managed clouds.

If neither card is a clean fit, look at the cards in between: the 32 GB RTX 5090, the 48 GB L40S and the 80 GB A100 all sit between these two on price and capability. The side-by-side comparison shows the two headline cards' specs and prices together.

From consumer to datacenter: VRAM and live rental prices
GPU VRAM Cheapest now Where Median across providers Providers
Nvidia GeForce RTX 4090 24GB $0.30/hr io.net $0.44/hr 10
Nvidia GeForce RTX 5090 32GB $0.45/hr Nosana $0.72/hr 10
Nvidia L40S 48GB $0.53/hr Lium $0.96/hr 12
Nvidia A100 80GB PCIe 80GB $0.65/hr Lium $1.51/hr 9
Nvidia H100 80GB $2.29/hr Theta EdgeCloud $2.81/hr 3
Nvidia H200 141GB $2.95/hr Vast.ai $3.69/hr 5

Live data: median hourly price per GPU, collected in the last 24 hours. How we collect prices.

Live provider comparison

The cheapest listings for each card by provider, from the last 24 hours of data. Marketplace listings vary from host to host, so check the offer count as well as the price.

Nvidia GeForce RTX 4090 — live prices by provider
ProviderMedian $/hrRangeOffers
io.net $0.30/hr $0.30/hr–$0.30/hr 792
Novita AI $0.33/hr $0.33/hr–$0.33/hr 11
RunPod $0.34/hr $0.34/hr–$0.34/hr 1
Nosana $0.36/hr $0.29/hr–$0.50/hr 6

Live data, last 24 hours. Full Nvidia GeForce RTX 4090 page.

Nvidia H100 — live prices by provider
ProviderMedian $/hrRangeOffers
Theta EdgeCloud $2.29/hr $2.29/hr–$2.29/hr 1
Fluence $2.81/hr $2.71/hr–$3.04/hr 3
Thunder Compute $3.20/hr $3.20/hr–$3.20/hr 0

Live data, last 24 hours. Full Nvidia H100 page.

FAQ

Is the RTX 4090 really cheaper per FLOP than the H100?
Usually, yes. On spec sheets the H100 SXM delivers about 989 dense BF16 tensor TFLOPS against roughly 165 for the RTX 4090 (with FP32 accumulation) — about 6× more — but it typically rents for more than 6× the price of a 4090 on marketplaces. The article computes the cost per petaflop-hour from today's live prices.
What's the catch with renting an RTX 4090 instead of an H100?
Memory and scaling. 24 GB holds about a 7–8B model at FP16 for inference, and far less for full training, while the H100 has 80 GB. And 4090s have no NVLink, so multi-GPU jobs communicate over PCIe and scale much worse than an NVLink-connected H100 node.
Which providers rent both cards?
Marketplaces and platforms such as Vast.ai and RunPod commonly list both; operator clouds like Lambda and the hyperscalers focus on datacenter GPUs. The live tables in the article show who lists each card right now.
When should I just pay for H100s?
For full fine-tuning or training beyond a few billion parameters, for jobs that need more than 24 GB per GPU, for multi-GPU training that depends on fast interconnect, and for production serving where predictable latency matters. The 4090 wins for LoRA/QLoRA fine-tuning of small models, batch inference and image generation.
Related