H100 vs H200 vs B200 for LLM inference.
Memory, bandwidth and cost per token for the H100, H200 and B200: when extra HBM avoids sharding, and which card to rent for your model size.
By Vi Nguyen · Published · Prices on this page are live · Editorial policy
If you are serving a large language model in production, the choice between an H100, an H200 and a B200 is rarely about peak FLOPS. For inference, and especially for the token-by-token decode phase that dominates chat and agent workloads, two numbers decide almost everything: how much memory the card has and how fast it can read that memory. This guide walks through both, shows how to turn an hourly rental price into a cost per million tokens, and ends with a decision framework keyed to model size.
The spec sheet numbers that matter for inference
All three are NVIDIA datacenter parts, but they span two architectures. The H100 and H200 are both Hopper: same compute die, same tensor cores, same FP8 support. The H200 is essentially an H100 with a much bigger and faster memory system bolted on. The B200 is Blackwell, a new generation with more compute, a faster NVLink, native FP4 support in its tensor cores, and more memory again.
| GPU (SXM form factor) | Architecture | Memory | Memory type | Bandwidth | Lowest native precision |
|---|---|---|---|---|---|
| H100 SXM | Hopper | 80 GB | HBM3 | 3.35 TB/s | FP8 |
| H200 SXM | Hopper | 141 GB | HBM3e | 4.8 TB/s | FP8 |
| B200 | Blackwell | 192 GB | HBM3e | 8 TB/s | FP4 |
Two caveats before you use that table to compare listings. First, the name on a rental listing does not always tell you which variant you are getting. The H100 PCIe uses HBM2e and has roughly 2 TB/s of bandwidth, well under the SXM part, and it runs at a lower power limit. The H100 NVL is a PCIe card designed to be bridged in pairs, with its own memory configuration and power limit, so it is not interchangeable with either. The H200 NVL is likewise the PCIe, air-cooled-server version of the H200, with the same 141 GB but a lower power envelope than the H200 SXM. When a provider just says "H100", ask or check the instance details: a PCIe card at an SXM price is a bad deal for bandwidth-bound work.
Second, the B200's 192 GB is the full-chip figure. Some HGX B200 systems expose a little less usable memory per GPU, so plan capacity with a margin rather than to the last gigabyte.
Why memory bandwidth sets your decode speed
LLM inference has two phases. Prefill processes your whole prompt in one pass; it is a large matrix multiply and is compute-bound. Decode generates output one token at a time, and for every single token the GPU has to stream the model's weights out of memory and through the tensor cores. At small batch sizes the tensor cores spend most of their time waiting for memory. That is why decode throughput tracks memory bandwidth far more closely than it tracks TFLOPS.
You can derive a rough ceiling from first principles. For one sequence at batch size 1:
max tokens/sec ≈ memory bandwidth ÷ bytes of weights read per token
Take a dense 70B model quantized to FP8, so about 70 GB of weights (1 byte per parameter). The estimated single-stream ceilings are:
- H100 SXM: 3,350 GB/s ÷ 70 GB ≈ 48 tokens/sec
- H200 SXM: 4,800 GB/s ÷ 70 GB ≈ 69 tokens/sec
- B200: 8,000 GB/s ÷ 70 GB ≈ 114 tokens/sec
These are upper bounds, not measurements. Real systems land below them because of kernel overhead, attention over the KV cache, and imperfect bandwidth utilization. But the ratios hold up well: the H200 has about 1.4x the H100's bandwidth, and the B200 about 2.4x. If your workload is interactive chat where each user cares about their own token stream, that ratio is roughly the speedup you should expect per stream.
Batching changes the economics. When you serve 32 requests at once, the weights are read once per step and used for all 32 sequences, so aggregate throughput climbs steeply until you either run out of compute or run out of memory for the KV cache. That second limit is where the H200 and B200 earn their price.
When extra memory saves you from sharding
Memory capacity determines two things: whether the model fits on one GPU at all, and how many concurrent sequences you can hold in KV cache once it does. Use the site's rule of thumb: weights in GB ≈ parameters in billions × bytes per parameter (2 for FP16/BF16, 1 for FP8, about 0.5 for INT4), then add the KV cache and 10 to 20 percent for runtime overhead. The VRAM fit calculator does this for specific models.
Worked example: Llama 3.3 70B at FP8
- Weights: 70B × 1 byte = 70 GB.
- KV cache per token: Llama 3 70B has 80 layers, 8 KV heads and a head dimension of 128. At FP16 that is 2 (K and V) × 80 × 8 × 128 × 2 bytes ≈ 0.33 MB per token.
- A single 32K-token conversation therefore needs about 32,768 × 0.33 MB ≈ 10.7 GB of KV cache.
- Runtime overhead (CUDA context, activations, allocator slack): budget about 8 to 10 GB.
On an 80 GB H100, 70 GB of weights plus overhead leaves almost nothing for KV cache. You can run it, but only with short contexts and tiny batches, which wrecks cost per token. In practice people run 70B on two H100s with tensor parallelism. On a 141 GB H200 the same model leaves around 60 GB for KV cache, enough for roughly five or six full 32K conversations or dozens of shorter ones, on a single card with no inter-GPU communication. On a B200 you get over 100 GB of headroom.
Sharding is not free. Tensor parallelism splits each layer across GPUs and adds an all-reduce on every layer of every token. Over NVLink inside an 8-GPU SXM node that cost is modest, and it lets you pool both bandwidth and capacity. Over PCIe, between separate rental instances, or across mixed hardware it can eat a large share of your throughput. Every GPU you avoid also removes a failure point and simplifies deployment. The rule: if a larger card lets you drop from two GPUs to one, or from 16 to 8, it is usually worth a higher hourly rate.
Live prices for the three generations
Rental prices for these cards move daily and vary a lot between marketplaces and managed clouds, so the table below is pulled from current listings rather than typed in. Pay attention to the variant column names: SXM and PCIe listings are tracked separately where providers label them.
| GPU | VRAM | Cheapest now | Where | Median across providers | Providers |
|---|---|---|---|---|---|
| Nvidia H100 | 80GB | $2.29/hr | Theta EdgeCloud | $2.81/hr | 3 |
| Nvidia H100 SXM | 80GB | $1.30/hr | Lium | $3.20/hr | 13 |
| Nvidia H200 | 141GB | $2.63/hr | Vast.ai | $3.61/hr | 7 |
| Nvidia H200 SXM | 141GB | $3.59/hr | RunPod | $3.99/hr | 2 |
| Nvidia B200 | 192GB | $3.69/hr | DeepInfra | $5.98/hr | 7 |
Live data: median hourly price per GPU, collected in the last 24 hours. How we collect prices.
Right now the cheapest H200 listing costs about 1.15x the cheapest H100 listing. Compare that ratio to the 1.43x bandwidth ratio and the 1.76x memory ratio between the two cards: if the price premium is below the bandwidth ratio, the H200 is cheaper per token for bandwidth-bound decode even before you count the batching headroom.
For the H100 specifically, the spread between providers is wide enough that it is worth looking at the per-provider breakdown:
| Provider | Median $/hr | Range | Offers |
|---|---|---|---|
| Theta EdgeCloud | $2.29/hr | $2.29/hr–$2.29/hr | 1 |
| Fluence | $2.81/hr | $2.71/hr–$3.04/hr | 3 |
| Thunder Compute | $3.20/hr | $3.20/hr–$3.20/hr | 0 |
Live data, last 24 hours. Full Nvidia H100 page.
The cheapest rows usually come from host-supplied marketplaces such as Vast.ai, where reliability and interconnect vary by host. Managed clouds such as Lambda cost more but give you consistent hardware and full 8-GPU nodes. RunPod sits in between with its Community and Secure clouds. Our Lambda vs RunPod vs Vast.ai comparison covers that trade-off in depth.
Turning an hourly price into cost per token
Hourly price is the wrong unit for comparing inference hardware. What you actually buy is tokens. The formula is:
$ per 1M tokens = hourly price ÷ (aggregate tokens/sec × 3,600) × 1,000,000
"Aggregate" means total output tokens per second across all concurrent requests on that GPU, at the batch size you actually run in production. Here is a worked example with round hypothetical numbers, chosen only to show the mechanics:
- Suppose an H100 rents at a hypothetical $2.50/hr and your serving stack sustains 1,500 aggregate tokens/sec on your model at your batch size.
- Tokens per hour: 1,500 × 3,600 = 5.4M.
- Cost per 1M tokens: $2.50 ÷ 5.4 = $0.46.
- Now suppose an H200 rents at a hypothetical $3.50/hr, and because it holds more KV cache it sustains a larger batch at 2,400 tokens/sec.
- Tokens per hour: 2,400 × 3,600 = 8.64M. Cost per 1M tokens: $3.50 ÷ 8.64 = $0.41.
The pricier card wins on unit cost. That pattern is common for memory-hungry models and long contexts, and it reverses for small models that already fit comfortably with a big batch on an H100. The only way to know your throughput figure is to measure it on your own model, prompt mix and serving engine, for a few hours on a rented card, before committing. Plug the result into the self-host vs API calculator to see whether renting beats a hosted endpoint at all. Remember to count utilization: a GPU that sits at 30 percent load for most of the day costs three times as much per token as the formula suggests.
Decision framework by model size
Assuming FP8 weights and a production workload that needs real context lengths and batching:
- Up to about 30B parameters. Pick the H100, or even look below it; a 30B model at FP8 is about 30 GB and fits one H100 with lots of KV room. The H200's extra memory buys little here, and the 48 GB cards in our 48 GB GPU comparison may be cheaper still. Only move to H200 if per-stream latency matters and the premium is small.
- 70B-class dense models. This is the H200's home ground. One H200 serves a 70B at FP8 with room to batch, where the H100 needs two cards. If you are already committed to H100 nodes, two-way tensor parallelism over NVLink works well. See the cheapest way to run a 70B model for lower-cost options.
- 100B to 200B, or 70B with very long contexts. One B200 or two H200s. A 120B model at FP8 is 120 GB of weights; that does not leave enough KV room on a single H200 but fits on a B200.
- 400B dense or large mixture-of-experts models. You need a full node. 405B at FP8 is about 405 GB, which fits 8× H100 (640 GB) but leaves limited cache, fits 4× or 8× H200 more comfortably, and is easiest on B200 nodes. For DeepSeek-class models see which GPUs can run DeepSeek and Qwen MoE models.
- Interactive latency is the product. If time-per-output-token is what users feel, the B200's bandwidth is the most direct lever you can buy, and its FP4 path roughly halves bytes per token again on models that tolerate it.
- Batch jobs with flexible deadlines. Take whichever card gives the lowest cost per token on your measured throughput, often a cheap marketplace H100, and accept some interruption risk.
Practical gotchas when renting these cards
Availability is uneven. H100s are now widely listed across marketplaces and clouds. H200s are less common, and B200s are mostly available as full 8-GPU nodes from managed providers, sometimes with minimum commitments. If you only need one GPU, a B200 may simply not be rentable in that shape.
Software support matters too. FP8 is mature on Hopper in vLLM, SGLang and TensorRT-LLM. Blackwell's FP4 kernels are newer, so check that your serving engine and quantization format support the card before assuming the headline speedup. Finally, check the node topology: eight SXM GPUs on NVSwitch behave very differently from eight PCIe cards in one chassis, and very differently again from eight single-GPU instances. For more on the costs that do not show up in the hourly rate, read the hidden costs of renting GPUs, and for the capacity math in more detail, how much VRAM you need to run LLMs.
The short version: the H100 is the default and the cheapest per hour, the H200 is often cheaper per token for 70B-class models because it avoids sharding and batches deeper, and the B200 is the card to rent when per-stream speed or very large models justify a full Blackwell node.