- Can I run it locally?
- gpt-oss
gpt-oss VRAM requirements: can I run gpt-oss-20b or 120b locally?
gpt-oss-20b fits a 16 GB card with room for context, and gpt-oss-120b wants about 67.5 GB. The quant you pick barely changes either number.
Check your machine
Fits on the card 13.4 GB
Everything stays in video memory, which is the fast case.
- Weights 11.6 GB
- Context cache 0.8 GB
- Overhead 0.9 GB
- Fast memory
- With system RAM
What we found
Here's the odd thing about gpt-oss. gpt-oss-20b's F16 file is 13.8 GB and its smallest, Q3_K_S, is 11.5 GB. OpenAI released the expert weights already in a 4-bit format called MXFP4, and the conversions leave them as they are, so going from the biggest file to the smallest saves about 2 GB. I'd take the F16 file.
At F16 with a 32K context the 20b comes to 15.5 GB, which a 16 GB card holds. Half its layers use a 128-token sliding window, so even its full 128K context only adds 3.2 GB of cache.
gpt-oss-120b is 67.5 GB the same way. That's a 96 GB Mac, a DGX Spark or an RTX PRO 6000. It runs four experts per token out of 128, so a 24 GB card with 64 GB of system RAM can run it too, with the experts left in RAM.
gpt-oss-20b
20.9 billion parameters · mixture of experts, 4 of 32 active · up to 128K context · 520 thousand downloads in 30 days
gpt-oss-20b needs about 13.4 GB at Q4_K_M with a 32K-token context. The file itself is 11.6 GB, and the cache for 32K tokens adds 0.8 GB. That fits a 16 GB card (RTX 4060 Ti 16 GB, 5070 Ti, 5080). On a Mac, 24 GB of unified memory is enough. It routes each token through 4 of 32 experts, so it runs faster than its size suggests.
12 of its layers only look back 128 tokens, so they stop growing early, and the whole cache at 128K is 3.2 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
Q2_K | 11.5 GB | 12.6 GB | 13.2 GB | 16 GB | 24 GB |
Q3_K_M | 11.5 GB | 12.6 GB | 13.2 GB | 16 GB | 24 GB |
Q4_K_M | 11.6 GB | 12.8 GB | 13.4 GB | 16 GB | 24 GB |
Q5_K_M | 11.7 GB | 12.8 GB | 13.5 GB | 16 GB | 24 GB |
Q6_K | 12.0 GB | 13.2 GB | 13.8 GB | 16 GB | 24 GB |
Q8_0 | 12.1 GB | 13.2 GB | 13.8 GB | 16 GB | 24 GB |
F16 | 13.8 GB | 14.9 GB | 15.5 GB | 16 GB | 24 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
gpt-oss-120b
117 billion parameters · mixture of experts, 4 of 128 active · up to 128K context · 75 thousand downloads in 30 days
At Q4_K_M, the gpt-oss-120b file is 62.8 GB, and with 32K tokens of context and runtime overhead it comes to about 64.9 GB. The smallest card that holds it whole is a RTX PRO 6000 (96 GB). On a Mac, 96 GB of unified memory is enough. It's a mixture-of-experts model (4 of 128 experts per token), which is why people run it with the experts parked in system RAM and the rest on a 24 GB card.
A sliding window of 128 tokens on 18 layers keeps the cache small: 4.9 GB at the full 128K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
Q2_K | 62.6 GB | 63.8 GB | 64.7 GB | 96 GB | 96 GB |
Q3_K_M | 62.6 GB | 63.9 GB | 64.8 GB | 96 GB | 96 GB |
Q4_K_M | 62.8 GB | 64.0 GB | 64.9 GB | 96 GB | 96 GB |
Q5_K_M | 62.9 GB | 64.1 GB | 65.0 GB | 96 GB | 96 GB |
Q6_K | 63.3 GB | 64.5 GB | 65.4 GB | 96 GB | 96 GB |
Q8_0 | 63.4 GB | 64.6 GB | 65.5 GB | 96 GB | 96 GB |
F16 | 65.4 GB | 66.6 GB | 67.5 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.