Công cụ · Chọn GPU
Nên thuê GPU nào?
Trả lời ba câu hỏi để nhận ba GPU tốt nhất nên thuê cho công việc của bạn ngay lúc này — kèm nhà cung cấp, giá theo giờ và chi phí một tháng.
Lựa chọn của chúng tôi
Cần khoảng 36 GB bộ nhớ GPU (INT8, 8K context, 4 concurrent).
Xếp theo giá trên mỗi đơn vị băng thông bộ nhớ — tốc độ sinh cao nhất trên mỗi đô la.
Danh sách thiết lập
- Use a serving engine with continuous batching (vLLM, SGLang or TGI) — it's what makes concurrent requests cheap.
- Download weights in the quantized format you sized for (AWQ/GPTQ for 4-bit, FP8 on Hopper/Ada) rather than quantizing on the box.
- Put model files on a persistent volume so a restart doesn't re-download tens of GB.
- Load-test with your real prompt lengths before sending traffic; long contexts shrink how many requests fit in memory.
Ad
Cách tính
The wizard starts from memory, because a GPU that can't hold the job is useless at any price. For running a model we size the weights at the precision your priority implies — 4-bit for cheapest, 8-bit for balanced, 16-bit for fastest — plus KV cache for an 8K context. Fine-tuning uses QLoRA (or LoRA when you pick fastest), training uses full mixed-precision state, and image models use typical VRAM needs for each model family.
Among the GPUs that fit and have a live price, cheapest ranks by hourly rate, fastest by memory bandwidth (which sets generation speed), and balanced by price per unit of bandwidth — the most speed per dollar. If no single card fits, it shows the smallest multi-GPU shapes instead. Every result links to the GPU's page with all providers, because the cheapest provider isn't always the most reliable one.