Provider comparison

Lambda vs RunPod vs Vast.ai: When to pick which.

A marketplace, a two-tier platform and an operator cloud: live H100 and A100 prices on each, what the price gap buys, and which fits your workload.

By Vi Nguyen · Published · Updated · Prices on this page are live · Editorial policy

Lambda, RunPod and Vast.ai show up in almost every "where should I rent a GPU?" thread, and they're often compared as if they were three versions of the same product. They aren't. One is a peer-to-peer marketplace, one is a platform that runs both a marketplace-style tier and its own datacenter tier, and one is an operator that owns its fleet. The hourly price differences between them mostly follow from that, and so do the reliability, support and compliance differences. Pick the model first and the provider usually picks itself.

Live prices: the same GPU on all three

The tables below are live. They show the median hourly price each of the three currently lists for one GPU, so you can see the spread today rather than on the day this was written.

Nvidia H100 — live prices by provider

No provider has listed a Nvidia H100 in the last 24 hours.

Live data, last 24 hours. Full Nvidia H100 page.

Nvidia A100 80GB PCIe — live prices by provider

No provider has listed a Nvidia A100 80GB PCIe in the last 24 hours.

Live data, last 24 hours. Full Nvidia A100 80GB PCIe page.

If a provider is missing from a table, it had no listing for that exact GPU variant in the last 24 hours. H100s come in SXM, PCIe and NVL forms that perform differently, and providers don't all stock the same one; the H100 page lists every variant and provider.

Three different business models

Vast.ai — a peer-to-peer marketplace

Vast.ai doesn't own the GPUs you rent. Independent hosts — small datacenters, operators with spare capacity, former crypto miners — list machines, and Vast provides the marketplace, the container runtime and billing. Prices are set by competition between hosts, which is why Vast is so often the cheapest place to find a given card. The same mechanism explains the downside: two listings for the "same" GPU can differ in CPU, disk speed, network bandwidth, and how reliably the host keeps the machine online.

Good fit: experiments, batch inference, fine-tuning runs you checkpoint, image generation, and any workload where a host failure costs you a restart rather than a customer.

Poor fit: latency-sensitive production serving, data you're not comfortable placing on a third party's hardware, and multi-node training — separate hosts don't share a fast interconnect.

RunPod — a platform with two supply tiers

RunPod sells capacity through a Community Cloud (third-party hosts that RunPod vets) and a Secure Cloud (capacity in partner datacenters with stricter standards), plus a serverless product for autoscaling inference endpoints. Its main draw is ergonomics: ready-made templates for Jupyter, ComfyUI, vLLM and similar stacks mean you can go from sign-up to a running model without writing a Dockerfile. Community Cloud prices typically sit close to marketplace levels; Secure Cloud costs more.

Good fit: individuals and small teams who want low prices without managing host selection, and inference endpoints that should scale to zero when idle.

Poor fit: the very cheapest possible hour (a marketplace usually wins) and large reserved clusters.

Lambda — an operator with its own fleet

Lambda runs its own GPU cloud: on-demand instances plus reserved capacity, including multi-node clusters with high-speed interconnect for distributed training. You're renting from the company that operates the hardware, so instance shapes are consistent and support is the operator's own. List prices are higher than marketplace medians and change rarely, which also makes them easy to budget.

Good fit: multi-GPU and multi-node training, teams that need a predictable provider relationship, and organizations with procurement or compliance requirements. Check Lambda's current security and compliance documentation against your own requirements rather than assuming.

Poor fit: short, bursty jobs where you'd pay the premium for predictability you don't need.

What the price gap buys you

It's tempting to read the tables above and conclude the cheapest provider wins. For a lot of work, it does. But the gap between a marketplace median and an operator's list price is paying for specific things, and it's worth knowing which ones you actually need:

  • Consistency. On an operator cloud, two instances of the same type are the same machine. On a marketplace, you choose the host, and a poor choice shows up as slow data loading or a mid-run disappearance.
  • Interconnect. Multi-GPU training across nodes needs fast networking between them. That exists inside an operator's cluster; it doesn't exist between unrelated hosts.
  • Capacity guarantees. Reserved contracts give you the same GPUs next month. Marketplace supply fluctuates, and popular cards can be scarce at peak times.
  • Accountability. When something breaks on an operator cloud, one company is responsible for fixing it.

If none of those matter for a job, the premium is wasted. If one of them is critical, a lower hourly rate that fails you once can cost more than the difference.

A worked example: one month of a 24/7 inference box

Say you serve a quantized 70B model on one 80 GB GPU around the clock — about 730 hours a month. Using today's live medians for an A100 80GB:

  • Vast.ai: no current A100 80GB listing
  • RunPod: no current A100 80GB listing
  • Lambda: no current A100 80GB listing

Now add the costs the hourly rate hides. Downloading a 40 GB model onto a fresh host takes time you pay for. A persistent volume keeps billing while an instance is stopped. If a marketplace host drops out twice a month and it takes you an hour each time to notice and redeploy, that's lost service rather than lost dollars — which matters more depends on who's on the other end. Our guide to the hidden costs of renting GPUs goes through each of these.

Decision framework

  • Exploring, prototyping, or on a tight budget → Vast.ai. Filter for hosts with a high reliability score and a history of rentals, and run a short test before a long job.
  • Want low prices without choosing hosts → RunPod Community Cloud, starting from a template.
  • Inference endpoint with uneven traffic → a serverless option such as RunPod's, so you pay for requests rather than idle hours — or a hosted API (see self-host vs API breakeven).
  • Steady production serving → RunPod Secure Cloud or Lambda; the reliability is worth the premium.
  • Multi-node distributed training → Lambda (or a hyperscaler), where the interconnect exists.
  • Formal compliance requirements → shortlist operators and verify their current certifications directly; don't rely on a comparison article for this.

Before you commit to any of them

  1. Run a one-hour test on the exact instance type: time the model download, measure throughput with your own workload, and check disk and network speed.
  2. Read the billing granularity and what continues billing when an instance is stopped (storage almost always does).
  3. Check egress pricing if your workload sends a lot of data out.
  4. Script your setup so moving to another provider takes minutes. Being able to switch is what lets you take advantage of the price gaps in the tables above.

FAQ

Which provider is absolutely cheapest?
On raw hourly price, Vast.ai usually undercuts the others because independent hosts compete on price, and RunPod's Community Cloud is often close. Lambda's list prices are typically higher but stable. The live tables in the article show today's numbers for H100 and A100 80GB.
Is RunPod or Vast.ai better for hobbyists?
RunPod's templates (Jupyter, ComfyUI, vLLM and others) make setup fast and remove host selection. Vast.ai is usually cheaper but you choose the host yourself, so you'll spend more time checking reliability.
Why would I pay Lambda's premium?
For predictability: consistent instance shapes on hardware Lambda operates itself, reserved capacity for longer commitments, multi-node clusters with fast interconnect, and one company accountable when something breaks. If you have compliance requirements, verify Lambda's current certifications directly.
Can I run distributed training across these?
Multi-node training needs fast networking between machines, which operator clouds like Lambda offer in their clusters. Multi-GPU instances on a single machine are available on all three. On Vast.ai, separate hosts are unrelated machines, so training across hosts is impractical.
Related