---
title: "Best cloud GPU for LLM inference in 2026, by model size | PowerGPU"
description: "From 8B to 405B: the cheapest rentable card that holds each model, indicative tokens per second with vLLM, and the cost per million tokens on today's sheet."
url: https://powergpu.ai/guides/best-cloud-gpu-for-llm-inference
last_modified: 2026-09-14T11:04:20+00:00
prices_as_of: 2026-09-14
site: PowerGPU (powergpu.ai)
---

Guide · Choosing hardware

# Best cloud GPU for LLM inference in 2026, by model size

From 8B to 405B: the cheapest rentable card that holds each model, indicative tokens per second with vLLM, and the cost per million tokens on today's sheet.

10 min read Published 2026-09-03 Updated 2026-09-03 (prices live from the sheet)

TL;DR

- Match VRAM to the model, then buy bandwidth: RTX 5090 for 7-14B, L40S or A6000 for 32B, H100 or H200 for 70B and up.
- Budget about 2.4 GB per billion parameters in FP16 and 0.62 GB in 4-bit, then add KV-cache for context times concurrency.
- On PowerGPU an RTX 5090 at $0.439 sustains roughly 2,800 tokens per second on 8B; interruptible capacity halves the cost per token.
- Cost per million tokens is hourly rate divided by tokens per second times 3,600, times one million; measure it at real concurrency.

## Three rules before any benchmark

1. **VRAM decides what runs.** Weights plus KV-cache must fit, with headroom: ~2.4 GB per billion parameters in FP16, ~0.62 GB in 4-bit, then add the cache for your context × concurrency. The [VRAM guide](https://powergpu.ai/guides/llm-vram-requirements) has the full tables.
2. **Bandwidth decides how fast.** Decoding is memory-bound: tokens per second track GB/s more than TFLOPS. That is why a 1.79 TB/s RTX 5090 beats a 1 TB/s RTX 4090 by more than its FLOPS suggest, and why HBM cards dominate 70B serving.
3. **Batch decides the bill.** Continuous batching (vLLM, SGLang, TGI) turns one card into a dozen concurrent streams. Cost per million tokens is only meaningful at the concurrency you will actually run.

## The table: model size → card → $ per million tokens

Indicative aggregate throughput under vLLM continuous batching (roughly 1k tokens in, 1k out, batch 32–64) and the resulting cost at today's on-demand rates — snapshot 2026-09-14. Treat the tokens/s column as order-of-magnitude community numbers and benchmark your own model; the price column is exact.

| Model | Precision / weights | Card | On-demand | ~tokens/s | $ / M tokens | Note |
| --- | --- | --- | --- | --- | --- | --- |
| Llama 3.1 8B | FP16, 16 GB | [RTX 4090](https://powergpu.ai/gpu/rtx-4090) (24 GB) | $0.327 | ~2,200 | **$0.041** | The default: fits with KV-cache room on 24 GB. |
| Llama 3.1 8B | FP16, 16 GB | [RTX 5090](https://powergpu.ai/gpu/rtx-5090) (32 GB) | $0.439 | ~2,800 | **$0.044** | 78% more bandwidth than the 4090 shows up directly in tokens/s. |
| Llama 3.1 8B | 4-bit, ~6 GB | [RTX 3060](https://powergpu.ai/gpu/rtx-3060) (12 GB) | $0.042 | ~900 | **$0.013** | The cheapest working endpoint on the sheet. |
| Qwen 2.5 14B | FP16, 28 GB | [RTX 5090](https://powergpu.ai/gpu/rtx-5090) (32 GB) | $0.439 | ~1,500 | **$0.081** | 32 GB holds FP16 14B with modest context. |
| Qwen 2.5 32B | 4-bit, ~20 GB | [RTX 4090](https://powergpu.ai/gpu/rtx-4090) (24 GB) | $0.327 | ~700 | **$0.130** | AWQ/GPTQ 4-bit on 24 GB; quality within a point of FP16. |
| Qwen 2.5 32B | FP16, 64 GB | [H100 PCIE](https://powergpu.ai/gpu/h100-pcie) (80 GB) | $1.867 | ~2,400 | **$0.216** | FP8 on Hopper roughly doubles this again. |
| Llama 3.1 70B | 4-bit, ~40 GB | [RTX A6000](https://powergpu.ai/gpu/rtx-a6000) (48 GB) | $0.281 | ~350 | **$0.223** | Cheapest single card that holds a 70B. |
| Llama 3.1 70B | 4-bit, ~40 GB | [H100 PCIE](https://powergpu.ai/gpu/h100-pcie) (80 GB) | $1.867 | ~1,200 | **$0.432** | One 80 GB card, real batch sizes, FP8 KV-cache. |
| Llama 3.1 70B | FP8, 70 GB | [H200](https://powergpu.ai/gpu/h200) (141 GB) | $2.791 | ~2,000 | **$0.388** | 141 GB leaves room for long contexts and big batches. |
| Llama 3.1 405B | FP8, ~405 GB | [H200](https://powergpu.ai/gpu/h200) (141 GB) | $2.791 | ~600 | **$1.292** | Tensor-parallel across 4× H200 — figure is per GPU. |

Interruptible capacity halves every number in the price column. Formula, if you want to redo it with your own measurement: $/M = $/hr ÷ (tok/s × 3600) × 1,000,000.

## 7B–14B: consumer cards win

Below 16 GB of weights nothing beats GeForce silicon per dollar. The [RTX 5090](https://powergpu.ai/gpu/rtx-5090) ($0.439/hr, 32 GB GDDR7) is the current king of tokens per dollar for 8B FP16 and 14B models; the [RTX 4090](https://powergpu.ai/gpu/rtx-4090) ($0.327/hr) is a hair slower and cheaper; the [RTX 3090](https://powergpu.ai/gpu/rtx-3090) ($0.108/hr) is the budget pick when throughput matters less than the hourly rate. For always-on endpoints that must survive months of 24/7 duty, the datacenter [L4](https://powergpu.ai/gpu/l4) ($0.225/hr, 72 W) and [L40S](https://powergpu.ai/gpu/l40s) trade a little speed for passive cooling and ECC.

## 24B–32B: the 48 GB middle

A 32B model is ~64 GB in FP16 and ~20 GB in 4-bit. Quantized, it runs on a 24 GB card with short contexts; for serious concurrency you want 48 GB. The [L40S](https://powergpu.ai/gpu/l40s) ($0.514/hr) is the datacenter answer — FP8, 864 GB/s, built for 24/7 — and the [RTX A6000](https://powergpu.ai/gpu/rtx-a6000) ($0.281/hr) the cheapest 48 GB on the sheet. Compare them in [RTX 5090 vs L40S](https://powergpu.ai/compare/rtx-5090-vs-l40s) and [RTX 6000 Ada vs A6000](https://powergpu.ai/compare/rtx-6000ada-vs-rtx-a6000).

## 70B and up: HBM or nothing

Seventy billion parameters is ~40 GB in 4-bit and ~140 GB in FP16. One 80 GB card serves the quantized model with real batch sizes: the [H100 PCIe](https://powergpu.ai/gpu/h100-pcie) ($1.867/hr) with FP8 KV-cache, or the [A100 PCIe](https://powergpu.ai/gpu/a100-pcie) ($0.374/hr) when budget wins. Full-precision 70B, 100B+ models and long-context serving belong on the [H200](https://powergpu.ai/gpu/h200) ($2.791/hr, 141 GB, 4.8 TB/s) or tensor-parallel across 2–8× H100 SXM on one machine — the per-GPU price does not change with count. Frontier-scale inference with FP4 is where the [B200](https://powergpu.ai/gpu/b200) ($5.425/hr) earns its rate: several times the H100's tokens per second at under 3× the price.

## Serving patterns that change the bill

- **Quantize first.** AWQ/GPTQ 4-bit halves the card class you need and rarely hurts chat quality. FP8 on Hopper/Ada/Blackwell is nearly free quality-wise and 1.5–2× faster.
- **Reserve the baseline, burst serverless.** A [reserved](https://powergpu.ai/products/reserved) card at −35% carries steady traffic; [serverless](https://powergpu.ai/products/serverless) workers absorb spikes and scale to zero.
- **Mount the model library read-only.** New workers skip the 40 GB download; cold starts become load-to-VRAM.
- **Prefix caching for agents.** If prompts share a long system prefix, [SGLang](https://powergpu.ai/templates/sglang) can double effective throughput over plain vLLM.
- **Measure, then commit.** Ten minutes on two candidate cards costs cents; the [vLLM guide](https://powergpu.ai/guides/serve-llm-vllm) ships the exact benchmark command.

---

Put the numbers to work

Every price in this guide is our live rate — fixed, ≥30% under the market median, billed per second. Deploy the exact setup above from the [console](https://cloud.powergpu.ai/) in about 30 seconds, paid in crypto, no card and no KYC.

---

*About PowerGPU:* PowerGPU (powergpu.ai) is a cloud GPU rental service offering 80 NVIDIA GPU models — from the RTX A2000 at $0.024/hr to the B300 — at fixed prices set at least 30% below the public GPU marketplace median and re-checked weekly (H100 SXM: $1.428/hr on-demand). Billing is per second with no minimums; payment is crypto only (USDT, BTC, XMR, ETH, SOL, LTC, TRX) with no KYC. Instances run in Tier-III datacenters across 32 regions with a 99.9% uptime SLA and deploy in about 30 seconds from the web console (cloud.powergpu.ai) or the REST API.

Source: https://powergpu.ai/guides/best-cloud-gpu-for-llm-inference · Site index for AI assistants: https://powergpu.ai/llms.txt · Full content: https://powergpu.ai/llms-full.txt
