Comparison

RTX 5090 vs RTX 4090 for running LLMs.

32 GB GDDR7 vs 24 GB GDDR6X: what fits on each card, how bandwidth sets token speed, cost per token at live prices, and when a 3090 is still the pick.

By Vi Nguyen · Published · Prices on this page are live · Editorial policy

The RTX 4090 has been the default consumer card for running open-weights LLMs since it launched, and the RTX 5090 is the obvious upgrade. Whether you are buying one for a home server or renting one by the hour on a marketplace, the question is the same: is the 5090's extra memory and bandwidth worth what it costs? The answer depends almost entirely on which models you want to run and how fast you need them to answer. This guide covers what fits in 24 GB versus 32 GB, why the move from GDDR6X to GDDR7 matters more than the compute uplift, how to compare the two on cost per token, and when an older RTX 3090 is still the smarter pick.

The spec differences that matter for LLMs

For language model inference, three spec-sheet numbers explain most of the difference between these cards: memory capacity, memory bandwidth, and which low-precision formats the tensor cores support natively.

GPU Architecture VRAM Memory type Bandwidth Native low precision
RTX 3090Ampere24 GBGDDR6X936 GB/sINT8 (no FP8)
RTX 4090Ada Lovelace24 GBGDDR6X1,008 GB/sFP8
RTX 5090Blackwell32 GBGDDR71,792 GB/sFP8, FP4

The headline is bandwidth. The 5090's GDDR7 on a wider 512-bit bus gives it about 1.78x the 4090's memory bandwidth, a much bigger jump than the 4090 made over the 3090 (under 10 percent). Capacity goes up by a third, from 24 GB to 32 GB. Power goes up too: the 5090 is rated at 575 W against the 4090's 450 W, which matters if you are building a multi-GPU box at home and does not matter at all if you are renting.

You will also see RTX 4090 D and RTX 5090 D listings. These are versions NVIDIA made for the Chinese market to comply with US export rules. The 4090 D keeps 24 GB but ships with fewer CUDA cores than the standard 4090. The 5090 D keeps 32 GB, and has been reported to restrict some AI workloads at the driver or firmware level; details have varied between revisions, so treat it as a different product rather than a regional 5090. Because memory size and type are the same, memory-bound decode may behave close to the standard parts, but compute-bound work such as prompt processing and batched serving may not. If a D variant is noticeably cheaper, run your own workload on it for an hour before relying on it.

What fits in 24 GB versus 32 GB

Use the standard rule of thumb: weights in GB ≈ parameters in billions × bytes per parameter (2 for FP16, 1 for FP8 or INT8, about 0.5 for 4-bit), plus KV cache for your context, plus 10 to 20 percent runtime overhead. Real 4-bit formats such as GGUF Q4_K_M or AWQ land a little above 0.5 bytes per parameter because of scales and a few layers kept at higher precision.

Worked example: a 32B model at 4-bit with 16K context

  1. Weights: 32B × about 0.55 bytes ≈ 18 GB.
  2. KV cache: a Qwen-style 32B model with 64 layers, 8 KV heads and 128-dimension heads needs 2 × 64 × 8 × 128 × 2 bytes ≈ 0.26 MB per token at FP16. For 16,384 tokens that is about 4.3 GB.
  3. Runtime overhead: about 1.5 to 2 GB for the CUDA context and activations.
  4. Total: roughly 24 GB.

That is right at the 4090's ceiling. It works, but you will be trimming context or quantizing the KV cache to 8-bit to make room. On the 5090 the same setup leaves about 8 GB spare, which buys you either 32K+ context or several concurrent sequences. That pattern repeats across model sizes:

  • 7B to 8B at FP16 (about 16 GB): comfortable on both. The 5090 simply holds more context.
  • 14B at FP16 (about 28 GB): does not fit a 4090; fits a 5090 with limited context. At FP8 (about 14 GB) it fits both easily.
  • 24B to 27B at FP8 (24 to 27 GB): 5090 only. On a 4090 you drop to 4-bit.
  • 32B at 4-bit (about 18 GB): tight on 24 GB, comfortable on 32 GB, as above.
  • 30B-class MoE models such as Qwen 3.5 35B-A3B at 4-bit: similar footprint to a dense 32B, but they decode much faster because only about 3B parameters are active per token.
  • 70B at 4-bit (about 38 to 40 GB): does not fit on either single card. You need two 24 GB cards, a 48 GB card, or partial CPU offload, which is slow.

If you want to check a specific model, the VRAM fit calculator runs this math for you, and how much VRAM you need to run LLMs explains each term in more depth.

Bandwidth: why GDDR7 is the real upgrade

When a model generates text, each new token requires reading essentially all of the active weights from VRAM. With one user and a small batch, the GPU's compute units spend most of that time waiting on memory, so decode speed is bounded by bandwidth. A useful estimate is:

max tokens/sec ≈ memory bandwidth ÷ bytes of weights read per token

For that 32B model at about 18 GB of 4-bit weights, the estimated single-stream ceilings are:

  • RTX 3090: 936 ÷ 18 ≈ 52 tokens/sec
  • RTX 4090: 1,008 ÷ 18 ≈ 56 tokens/sec
  • RTX 5090: 1,792 ÷ 18 ≈ 100 tokens/sec

These are theoretical upper bounds derived from the spec sheet, not benchmark results; real throughput lands below them and depends on the inference engine, quantization kernels and context length. The ratios are the useful part. For single-user chat on a model that fits both cards, expect the 5090 to generate noticeably faster than the 4090, and expect the 4090 to feel only slightly faster than a 3090. The 4090's big compute advantage over the 3090 shows up in prompt processing (prefill), image generation and batched serving, not in single-stream token generation.

The 5090's FP4 support is a second, softer advantage. Blackwell tensor cores run 4-bit floating-point formats natively, which can speed up compute-bound phases for models quantized that way. Tooling support for FP4 in consumer inference stacks is still maturing, so do not count on it until your engine of choice supports it for your model.

Rental prices and the price ratio

Because marketplace prices move daily, here are the current numbers pulled from live listings rather than typed in:

Consumer GPUs for local LLMs, cheapest and median hourly price right now
GPU VRAM Cheapest now Where Median across providers Providers
Nvidia GeForce RTX 5090 32GB $0.42/hr Vast.ai $0.73/hr 10
Nvidia GeForce RTX 4090 24GB $0.30/hr io.net $0.45/hr 11
Nvidia GeForce RTX 3090 24GB $0.12/hr Vast.ai $0.21/hr 8
Nvidia RTX 4090 D 24GB $0.27/hr Clore.ai $0.30/hr 2
Nvidia RTX 5090 D 32GB $0.41/hr Clore.ai $0.41/hr 1

Live data: median hourly price per GPU, collected in the last 24 hours. How we collect prices.

At the moment the cheapest 5090 listing is about 1.39x the price of the cheapest 4090 listing. Hold that number against the 1.78x bandwidth ratio: if the price ratio is lower, the 5090 is cheaper per generated token on bandwidth-bound work, before you even count the extra 8 GB.

Most of these listings come from host-supplied marketplaces such as Vast.ai and Clore.ai, where reliability, CPU, disk speed and network vary host to host. RunPod lists consumer cards in both its Community and Secure clouds. The cheapest row is not always the right one; check the host's reliability score and upload bandwidth, since pulling a 20 GB model over a slow link wastes paid hours. Our guide to the hidden costs of renting GPUs covers what else to check.

Cost per token: a worked comparison

To compare cards fairly, convert the hourly price to cost per million generated tokens:

$ per 1M tokens = hourly price ÷ (tokens/sec × 3,600) × 1,000,000

A worked example with hypothetical prices and the estimated ceilings above, for a single user on the 32B model:

  1. Suppose a 4090 rents at a hypothetical $0.40/hr and delivers 45 tokens/sec in practice (about 80 percent of the 56 tokens/sec ceiling). Tokens per hour: 45 × 3,600 = 162,000. Cost: $0.40 ÷ 0.162 = $2.47 per 1M tokens.
  2. Suppose a 5090 rents at a hypothetical $0.70/hr and delivers 80 tokens/sec (the same 80 percent of its 100 tokens/sec ceiling). Tokens per hour: 80 × 3,600 = 288,000. Cost: $0.70 ÷ 0.288 = $2.43 per 1M tokens.

At a 1.75x price ratio the two cards cost almost the same per token, and the 5090 gives you the answer nearly twice as fast. Below that ratio the 5090 wins outright; well above it the 4090 wins on cost while losing on latency. Batching changes the numbers a lot: serve many users at once and aggregate throughput rises well above these single-stream figures, and the 5090's extra 8 GB of KV cache headroom lets it batch deeper. That single-user figure is also far above what hosted APIs charge for similar models, which is why the self-host vs API calculator is worth running before you rent anything for a low-traffic app.

When the RTX 3090 is still the value pick

The 3090 has the same 24 GB as the 4090 and about 93 percent of its bandwidth. For single-stream decode of a model that fits, that makes it nearly as fast in tokens per second, and it typically rents for less. It is the value pick when:

  • You are running a single user or a small batch. Chat, coding assistants and agents on a personal endpoint are bandwidth-bound, where the 3090 nearly matches the 4090.
  • Your model fits in 24 GB at INT4 or INT8. You lose nothing on capacity compared with the 4090.
  • You want two cards for a 70B model. Two 3090s give you 48 GB for 70B at 4-bit, and the 3090 supports an NVLink bridge, which the 4090 and 5090 do not. Tensor parallelism over PCIe still works on the newer cards, but the pair of 3090s is often the cheapest 48 GB you can rent.

Skip it when you need FP8 (Ampere has no FP8 tensor cores), heavy prompt processing with long documents, image or video generation, or high-concurrency serving. In those compute-bound cases the 4090's roughly doubled tensor throughput pays for itself. For more sub-$1 options, see best GPUs under $1/hr for AI inference.

Decision framework

  • Your model is 8B to 14B at FP8 or 4-bit, one user: rent a 3090 or 4090, whichever is cheaper at the time. The 5090 is faster but the model does not need its memory.
  • Your model is a 24B to 32B dense model and you want long context: the 5090. On 24 GB you will be fighting for every gigabyte of KV cache.
  • Latency is the product (interactive coding, voice, agents with many steps): the 5090, because bandwidth converts directly into faster tokens.
  • Serving many users on a small model: compare 4090 and 5090 on measured cost per token at your batch size. The 5090 often wins because it batches deeper.
  • You need 70B: neither card alone. Rent two 24 GB cards, or look at single 48 GB cards in our 48 GB GPU comparison and the options in the cheapest way to run a 70B model.
  • The price gap is unusually wide or narrow today: let the live price ratio above decide. Below about 1.7x, lean 5090; above about 2x, lean 4090.

FAQ

Is the RTX 5090 worth it over the 4090 for LLMs?
For token generation it has about 1.78x the memory bandwidth and 8 GB more VRAM, so it is faster per stream and fits larger models or longer context. It is worth it when the rental price ratio is below the bandwidth ratio or when you need more than 24 GB.
What models fit in 32 GB but not 24 GB?
Dense 24B to 27B models at FP8, 14B at FP16, and 32B at 4-bit with long context. A 70B model at 4-bit needs about 40 GB and fits on neither card alone.
What are the RTX 4090 D and RTX 5090 D?
They are versions made for the Chinese market to comply with US export rules. The 4090 D has fewer CUDA cores than a standard 4090, and the 5090 D has been reported to restrict some AI workloads, so test your own workload before relying on one.
Is the RTX 3090 still good for running LLMs?
Yes for single-user inference of models that fit in 24 GB. It has about 93 percent of the 4090's memory bandwidth, so token generation is nearly as fast, and it supports an NVLink bridge for 48 GB pairs. It lacks FP8 and is much slower for prompt processing and image generation.
Can I run a 70B model on an RTX 5090?
Not on one card at useful precision. At 4-bit the weights are about 38 to 40 GB, more than 32 GB. Use two 24 GB or 32 GB cards, a 48 GB card, or a hosted API.
Related