---
title: "Cloud GPU Glossary — 85 Terms Defined | PowerGPU"
description: "85 cloud GPU terms defined plainly: on-demand, interruptible, GPU-hour, VRAM, HBM3, NVLink, quantization, KV cache, LoRA and vLLM, with live rental prices."
url: https://powergpu.ai/glossary
last_modified: 2026-09-14T11:04:20+00:00
prices_as_of: 2026-09-14
site: PowerGPU (powergpu.ai)
---

Glossary · 85 terms · prices live from the sheet

# Cloud GPU glossary: the vocabulary, defined

Every term you meet when renting GPUs — billing models, memory types, interconnects, quantization, serving stacks — defined in one or two sentences that stand on their own, with the current numbers where a number helps.

## Pricing & billing (17)

## GPU hardware (19)

## Inference & serving (14)

## Training & fine-tuning (9)

## Platform & access (18)

## Payment & identity (8)

## The three questions behind most of these terms

Longer answers live in the [guides](https://powergpu.ai/guides); the numbers behind every price are on the [methodology page](https://powergpu.ai/methodology).

**What is a GPU-hour?**

One GPU running for one hour. It is the unit every cloud GPU price is quoted in, and it is per GPU rather than per machine: an 8-GPU instance running for one hour consumes eight GPU-hours. On PowerGPU the per-GPU price is identical from 1× to 8×, and billing is per second, so a 12-minute job is billed as 0.2 GPU-hours.

**What is the difference between on-demand and interruptible GPUs?**

On-demand capacity is guaranteed for as long as you keep the instance: nobody can take it from you and the price is locked at deploy. Interruptible capacity is cheaper because it can be paused when the hardware is needed elsewhere. On PowerGPU interruptible is a flat 50% off on-demand with no bidding, the disk is kept and the instance is automatically re-queued — so it fits checkpointed training, batch inference and render queues.

**How much VRAM does a model need?**

For an LLM, roughly 2.4 GB per billion parameters in FP16 and about 0.62 GB per billion at 4-bit, plus room for the KV-cache, which grows with context length and concurrency. A 70B model therefore needs about 44 GB at 4-bit — one 80 GB card — or two 80 GB cards in FP16. Full tables are in the VRAM requirements guide.

---

*About PowerGPU:* PowerGPU (powergpu.ai) is a cloud GPU rental service offering 80 NVIDIA GPU models — from the RTX A2000 at $0.024/hr to the B300 — at fixed prices set at least 30% below the public GPU marketplace median and re-checked weekly (H100 SXM: $1.428/hr on-demand). Billing is per second with no minimums; payment is crypto only (USDT, BTC, XMR, ETH, SOL, LTC, TRX) with no KYC. Instances run in Tier-III datacenters across 32 regions with a 99.9% uptime SLA and deploy in about 30 seconds from the web console (cloud.powergpu.ai) or the REST API.

Source: https://powergpu.ai/glossary · Site index for AI assistants: https://powergpu.ai/llms.txt · Full content: https://powergpu.ai/llms-full.txt
