Best GPUs under $1/hr for AI inference.
Every GPU renting under $1/hr right now, sorted by VRAM — plus which models each can actually run, cost per token, and how to avoid unreliable hosts.
By Vi Nguyen · Published · Updated · Prices on this page are live · Editorial policy
A dollar an hour buys a surprising amount of inference in 2026 — if you pick for memory first and price second. Most of the value on this list is in cards that were never marketed for AI: consumer and workstation GPUs rented out by marketplace hosts. This guide lists what's under $1/hr right now, then explains which of those cards can actually run the model you have in mind, and how to avoid the cheap listings that cost you more in lost time than they save.
Live leaderboard: sub-$1/hr GPUs
Every GPU below has a current listing under $1/hr for a single GPU, with at least 8 GB of VRAM. It's sorted by memory, largest first, because memory decides what you can run and price only decides what you pay for it. The list refreshes continuously from provider listings; the price shown is the cheapest provider's current median.
| GPU | $/hr (cheapest) | VRAM | Provider | |
|---|---|---|---|---|
| $0.50/hr | 192GB | RunPod | Compare → | |
| $0.50/hr | 141GB | RunPod | Compare → | |
| $0.86/hr | 96GB | Clore.ai | Compare → | |
| $0.91/hr | 96GB | Clore.ai | Compare → | |
| $0.91/hr | 96GB | Nosana | Compare → | |
| $0.65/hr | 80GB | Lium | Compare → | |
| $0.99/hr | 80GB | Clore.ai | Compare → | |
| $0.32/hr | 48GB | Nosana | Compare → | |
| $0.32/hr | 48GB | Nosana | Compare → | |
| $0.53/hr | 48GB | Lium | Compare → | |
| $0.54/hr | 48GB | Vast.ai | Compare → | |
| $0.69/hr | 48GB | RunPod | Compare → | |
| $0.71/hr | 48GB | Clore.ai | Compare → | |
| $0.47/hr | 40GB | Vast.ai | Compare → | |
| $0.50/hr | 34GB | RunPod | Compare → | |
| $0.29/hr | 32GB | Theta EdgeCloud | Compare → | |
| $0.29/hr | 32GB | Clore.ai | Compare → | |
| $0.38/hr | 32GB | Vast.ai | Compare → | |
| $0.42/hr | 32GB | Clore.ai | Compare → | |
| $0.49/hr | 32GB | RunPod | Compare → | |
| $0.50/hr | 32GB | RunPod | Compare → | |
| $0.083/hr | 24GB | Clore.ai | Compare → | |
| $0.12/hr | 24GB | Vast.ai | Compare → | |
| $0.16/hr | 24GB | Lium | Compare → | |
| $0.18/hr | 24GB | Clore.ai | Compare → |
Start with memory, not price
For LLM inference the first question is whether the model's weights fit in the card's VRAM with room left for the KV cache (the per-conversation memory that grows with context length). The weights are easy to estimate: parameters × bytes per parameter. At FP16 that's 2 bytes, at INT8 1 byte, and at 4-bit quantization about half a byte.
| Model size | FP16 weights | INT8 | 4-bit | Smallest sensible card |
|---|---|---|---|---|
| 7–8B | ~15 GB | ~8 GB | ~4–5 GB | 12–16 GB (4-bit/INT8), 24 GB for FP16 |
| 13–14B | ~27 GB | ~14 GB | ~7–8 GB | 16 GB at 4-bit, 24 GB at INT8 |
| 32B | ~64 GB | ~32 GB | ~17–19 GB | 24 GB at 4-bit |
| 70B | ~140 GB | ~70 GB | ~35–40 GB | 48 GB at 4-bit, or two 24 GB cards |
Add roughly 10–20% for the runtime plus whatever the KV cache needs for your context length and number of concurrent requests. A 70B model at 4-bit does not fit on a single 24 GB card — it needs a 48 GB card or two 24 GB cards split with tensor parallelism. The VRAM fit calculator does this math for a specific model and context length, and our guide to how much VRAM an LLM needs explains the KV cache in detail.
The three groups of sub-$1 GPUs
24 GB consumer cards: the value sweet spot
The RTX 3090 and RTX 4090 both carry 24 GB, which covers 7–14B models at FP16 or INT8 and 32B-class models at 4-bit — the range most self-hosted assistants, coding models and RAG pipelines live in. The 4090 is considerably faster, especially for prompt processing and image generation; the 3090 usually rents for less. Right now the cheapest 3090 is $0.12/hr on Vast.ai and the cheapest 4090 is $0.30/hr on io.net. The newer RTX 5090 adds 32 GB and more bandwidth, and sometimes dips under $1 too.
16 GB and smaller cards: fine for 7–8B
Cards like the RTX 4070 Ti Super, RTX 4060 Ti and the datacenter T4 are the cheapest way to serve a single 7–8B model, an embedding model, or a speech-to-text model. They get tight quickly: long contexts or several concurrent users will exhaust 16 GB of KV cache headroom even when the weights fit. Check the exact VRAM of the listing — several consumer models ship in 8 GB and 16 GB versions.
48 GB workstation and datacenter cards: 70B territory
The interesting part of the list is at the top. 48 GB cards such as the RTX A6000, A40 and L40S regularly appear under a dollar on marketplaces, and they're the cheapest single cards that hold a 70B model at 4-bit. For a team that wants one private 70B endpoint, that's usually the best-value option on this page. See our comparison of 48 GB GPUs for the differences between them.
What a dollar an hour buys per token
For a steady workload, what matters is cost per million generated tokens: $/hr ÷ (tokens per second × 3,600) × 1,000,000. As an illustration: a card renting at a hypothetical $0.40/hr that sustains 400 tokens/s across batched requests works out to $0.40 ÷ 1,440,000 × 1,000,000 ≈ $0.28 per million tokens. The same card serving one user at a time at 40 tokens/s costs ten times as much per token. Utilization and batching matter more than the hourly price, which is why cheap GPUs only beat hosted APIs when they're kept busy — the self-host vs API calculator shows where that line is for your traffic.
Quantization: the other lever
4-bit formats such as AWQ, GPTQ and GGUF Q4 cut weight memory to roughly a quarter of FP16, which is what lets 32B models onto 24 GB cards and 70B models onto 48 GB cards. The quality cost depends on the model and the task: for chat and summarization it's often hard to notice, while math, code and long-context retrieval can degrade more. INT8 or FP8 is a safer middle ground when it fits. Whatever you choose, run your own evaluation prompts before and after quantizing rather than trusting a leaderboard.
The cheapest listing isn't always the cheapest option
Most sub-$0.50/hr listings come from marketplace hosts on Vast.ai, Clore.ai, Nosana and io.net. The GPU is the same silicon wherever you rent it; what varies is everything around it — CPU, disk and network speed, and how reliably the host keeps the machine up. Before trusting a cheap host with anything long-running:
- Prefer hosts with a visible reliability score and a history of rentals.
- Time the model download and a short benchmark on your own workload in the first 15 minutes, and move on if either is slow.
- Keep checkpoints and model files somewhere you control, so a host failure costs a restart, not a day.
If you need consistency more than the absolute lowest rate, a managed tier such as RunPod costs more per hour but removes host selection from your job. Our marketplaces vs managed clouds guide covers the trade-off.