Gemma VRAM requirements: can I run Gemma 4 locally?
Gemma 4 12B needs 8.9 GB at a 4-bit quant with 32K of context, which a 12 GB card holds. Here's every Gemma in wide use, from the phone-sized E2B to the 31B.
Check your machine
Fits on the card 8.9 GB
Everything stays in video memory, which is the fast case.
- Weights 7.1 GB
- Context cache 0.8 GB
- Overhead 1.0 GB
- Fast memory
- With system RAM
What we found
Gemma's cache hardly grows, and that's by design. In Gemma 4 five layers out of every six only look back over a sliding window of 1,024 tokens (512 on the E models), so only one layer in six keeps the whole conversation. Gemma 4 31B at its full 256K context carries 12.0 GB of cache. Gemma 3 27B needs 11.4 GB at half that context.
The E models go further. E4B's last 18 layers reuse keys and values computed by earlier layers, so they add nothing to the cache at all, and its whole 128K context costs 2.2 GB.
Gemma 4 12B is the most downloaded Gemma right now (1.2 million in 30 days) and the sensible pick for a 12 GB card. With 24 GB, Gemma 4 26B A4B fits at 18.6 GB, and since only a few of its experts run for each token it answers faster than the dense 31B.
Google also publishes QAT builds, trained so they lose less at 4-bit. They need the same memory as the regular file at the same quant, so they don't get separate sections here.
Gemma 4 12B
11.9 billion parameters · up to 256K context · 1.2 million downloads in 30 days
Gemma 4 12B needs about 8.9 GB at Q4_K_M with a 32K-token context. The file itself is 7.1 GB, and the cache for 32K tokens adds 0.8 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.
40 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 256K is 2.7 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 4.7 GB | 6.3 GB | 6.5 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 6.0 GB | 7.6 GB | 7.8 GB | 8 GB | 16 GB |
Q4_K_M | 7.1 GB | 8.7 GB | 8.9 GB | 12 GB | 16 GB |
Q5_K_M | 8.4 GB | 10.0 GB | 10.2 GB | 12 GB | 16 GB |
Q6_K | 9.8 GB | 11.4 GB | 11.6 GB | 12 GB | 16 GB* |
Q8_0 | 12.7 GB | 14.3 GB | 14.5 GB | 16 GB | 24 GB |
BF16 | 23.8 GB | 25.4 GB | 25.6 GB | 24 GB | 32 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Gemma 4 E4B
7.5 billion parameters · up to 128K context · 594 thousand downloads in 30 days
At Q4_K_M, the Gemma 4 E4B file is 5.0 GB, and with 32K tokens of context and runtime overhead it comes to about 6.6 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
A sliding window of 512 tokens on 20 layers keeps the cache small: 2.2 GB at the full 128K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 3.8 GB | 5.0 GB | 5.4 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 4.6 GB | 5.8 GB | 6.2 GB | 8 GB | 16 GB |
Q4_K_M | 5.0 GB | 6.2 GB | 6.6 GB | 8 GB | 16 GB |
Q5_K_M | 5.5 GB | 6.7 GB | 7.1 GB | 8 GB | 16 GB |
Q6_K | 7.1 GB | 8.3 GB | 8.7 GB | 12 GB | 16 GB |
Q8_0 | 8.2 GB | 9.4 GB | 9.8 GB | 12 GB | 16 GB |
BF16 | 15.1 GB | 16.3 GB | 16.7 GB | 16 GB | 24 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 4 26B A4B
25.2 billion parameters · mixture of experts, 128 experts · up to 256K context · 552 thousand downloads in 30 days
Gemma 4 26B A4B needs about 18.6 GB at UD-Q4_K_M with a 32K-token context. The file itself is 16.9 GB, and the cache for 32K tokens adds 0.7 GB. That fits a 24 GB card (RTX 3090, RTX 4090). It routes each token through a few of 128 experts, so it runs faster than its size suggests.
25 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 256K is 3.0 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 10.5 GB | 12.0 GB | 12.2 GB | 12 GB | 16 GB* |
UD-Q3_K_XL | 12.9 GB | 14.3 GB | 14.6 GB | 16 GB | 24 GB |
UD-Q4_K_M | 16.9 GB | 18.4 GB | 18.6 GB | 24 GB | 24 GB* |
UD-Q5_K_XL | 21.2 GB | 22.7 GB | 22.9 GB | 24 GB | 32 GB |
UD-Q6_K_XL | 23.3 GB | 24.7 GB | 25.0 GB | 24 GB | 32 GB* |
Q8_0 | 26.9 GB | 28.3 GB | 28.5 GB | 32 GB | 32 GB* |
BF16 | 50.5 GB | 51.9 GB | 52.2 GB | 96 GB | 64 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Gemma 4 31B
30.7 billion parameters · up to 256K context · 371 thousand downloads in 30 days
At Q4_K_M, the Gemma 4 31B file is 18.3 GB, and with 32K tokens of context and runtime overhead it comes to about 22.0 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090). On a Mac, 32 GB of unified memory is enough.
A sliding window of 1,024 tokens on 50 layers keeps the cache small: 12.0 GB at the full 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 11.8 GB | 14.4 GB | 15.4 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 15.4 GB | 18.0 GB | 19.0 GB | 24 GB | 24 GB* |
Q4_K_M | 18.3 GB | 21.0 GB | 22.0 GB | 24 GB | 32 GB |
Q5_K_M | 21.7 GB | 24.3 GB | 25.3 GB | 24 GB | 32 GB* |
Q6_K | 25.2 GB | 27.8 GB | 28.8 GB | 32 GB | 32 GB* |
Q8_0 | 32.6 GB | 35.3 GB | 36.3 GB | 48 GB | 48 GB |
BF16 | 61.4 GB | 64.0 GB | 65.1 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Gemma 3 27B
27.0 billion parameters · up to 128K context · 30 thousand downloads in 30 days
Gemma 3 27B needs about 20.9 GB at Q4_K_M with a 32K-token context. The file itself is 16.5 GB, and the cache for 32K tokens adds 3.3 GB. That fits a 24 GB card (RTX 3090, RTX 4090).
52 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 11.4 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 6.5 GB | 8.9 GB | 10.9 GB | 12 GB | 16 GB |
UD-Q2_K_XL | 10.7 GB | 13.0 GB | 15.1 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 13.7 GB | 16.1 GB | 18.1 GB | 24 GB | 24 GB* |
Q4_K_M | 16.5 GB | 18.9 GB | 20.9 GB | 24 GB | 24 GB* |
Q5_K_M | 19.3 GB | 21.6 GB | 23.6 GB | 24 GB | 32 GB* |
Q6_K | 22.2 GB | 24.5 GB | 26.5 GB | 32 GB | 32 GB* |
Q8_0 | 28.7 GB | 31.1 GB | 33.1 GB | 32 GB | 36 GB* |
BF16 | 54.0 GB | 56.4 GB | 58.4 GB | 96 GB | 64 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Gemma 4 E2B
4.6 billion parameters · up to 128K context · 452 thousand downloads in 30 days
At Q4_K_M, the Gemma 4 E2B file is 3.1 GB, and with 32K tokens of context and runtime overhead it comes to about 4.4 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
A sliding window of 512 tokens on 12 layers keeps the cache small: 0.8 GB at the full 128K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 2.4 GB | 3.5 GB | 3.7 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.9 GB | 4.0 GB | 4.2 GB | 8 GB | 16 GB |
Q4_K_M | 3.1 GB | 4.2 GB | 4.4 GB | 8 GB | 16 GB |
Q5_K_M | 3.4 GB | 4.5 GB | 4.6 GB | 8 GB | 16 GB |
Q6_K | 4.5 GB | 5.6 GB | 5.8 GB | 8 GB | 16 GB |
Q8_0 | 5.0 GB | 6.1 GB | 6.3 GB | 8 GB | 16 GB |
BF16 | 9.3 GB | 10.4 GB | 10.6 GB | 12 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 3 4B
3.9 billion parameters · up to 128K context · 84 thousand downloads in 30 days
Gemma 3 4B needs about 4.4 GB at Q4_K_M with a 32K-token context. The file itself is 2.5 GB, and the cache for 32K tokens adds 0.9 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
29 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 2.9 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 1.2 GB | 2.6 GB | 3.1 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 1.8 GB | 3.2 GB | 3.7 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.2 GB | 3.5 GB | 4.0 GB | 8 GB | 16 GB |
Q4_K_M | 2.5 GB | 3.9 GB | 4.4 GB | 8 GB | 16 GB |
Q5_K_M | 2.8 GB | 4.2 GB | 4.7 GB | 8 GB | 16 GB |
Q6_K | 3.2 GB | 4.6 GB | 5.1 GB | 8 GB | 16 GB |
Q8_0 | 4.1 GB | 5.5 GB | 6.0 GB | 8 GB | 16 GB |
BF16 | 7.8 GB | 9.2 GB | 9.7 GB | 12 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 3 270M
268 million parameters · up to 32K context · 63 thousand downloads in 30 days
At Q4_K_M, the Gemma 3 270M file is 0.3 GB, and with 32K tokens of context and runtime overhead it comes to about 1.4 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
A sliding window of 512 tokens on 15 layers keeps the cache small: 0.1 GB at the full 32K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 0.2 GB | 1.3 GB | 1.4 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 0.2 GB | 1.3 GB | 1.4 GB | 8 GB | 16 GB |
Q4_K_M | 0.3 GB | 1.3 GB | 1.4 GB | 8 GB | 16 GB |
Q5_K_M | 0.3 GB | 1.3 GB | 1.4 GB | 8 GB | 16 GB |
Q6_K | 0.3 GB | 1.4 GB | 1.4 GB | 8 GB | 16 GB |
Q8_0 | 0.3 GB | 1.4 GB | 1.4 GB | 8 GB | 16 GB |
F16 | 0.5 GB | 1.6 GB | 1.7 GB | 8 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 3 12B
11.8 billion parameters · up to 128K context · 62 thousand downloads in 30 days
Gemma 3 12B needs about 11.0 GB at Q4_K_M with a 32K-token context. The file itself is 7.3 GB, and the cache for 32K tokens adds 2.7 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.
40 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 9.1 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 3.1 GB | 5.1 GB | 6.7 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 4.9 GB | 6.9 GB | 8.5 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 6.1 GB | 8.2 GB | 9.8 GB | 12 GB | 16 GB |
Q4_K_M | 7.3 GB | 9.4 GB | 11.0 GB | 12 GB | 16 GB |
Q5_K_M | 8.4 GB | 10.5 GB | 12.1 GB | 12 GB | 16 GB* |
Q6_K | 9.7 GB | 11.7 GB | 13.3 GB | 16 GB | 24 GB |
Q8_0 | 12.5 GB | 14.6 GB | 16.2 GB | 16 GB | 24 GB |
BF16 | 23.5 GB | 25.6 GB | 27.2 GB | 32 GB | 32 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Gemma 3 1B
1000 million parameters · up to 32K context · 42 thousand downloads in 30 days
At Q4_K_M, the Gemma 3 1B file is 0.8 GB, and with 32K tokens of context and runtime overhead it comes to about 2.0 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
A sliding window of 512 tokens on 22 layers keeps the cache small: 0.2 GB at the full 32K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 0.6 GB | 1.7 GB | 1.8 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 0.7 GB | 1.8 GB | 1.9 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 0.7 GB | 1.8 GB | 1.9 GB | 8 GB | 16 GB |
Q4_K_M | 0.8 GB | 1.9 GB | 2.0 GB | 8 GB | 16 GB |
Q5_K_M | 0.9 GB | 1.9 GB | 2.0 GB | 8 GB | 16 GB |
Q6_K | 1.0 GB | 2.1 GB | 2.2 GB | 8 GB | 16 GB |
Q8_0 | 1.1 GB | 2.2 GB | 2.3 GB | 8 GB | 16 GB |
BF16 | 2.0 GB | 3.1 GB | 3.2 GB | 8 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 3n E2B
4.5 billion parameters · up to 32K context · 39 thousand downloads in 30 days
Gemma 3n E2B needs about 4.4 GB at Q4_K_M with a 32K-token context. The file itself is 3.0 GB, and the cache for 32K tokens adds 0.3 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
16 of its layers only look back 512 tokens, so they stop growing early, and the whole cache at 32K is 0.3 GB.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 2.6 GB | 3.8 GB | 4.0 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.8 GB | 4.0 GB | 4.2 GB | 8 GB | 16 GB |
Q4_K_M | 3.0 GB | 4.2 GB | 4.4 GB | 8 GB | 16 GB |
Q5_K_M | 3.3 GB | 4.4 GB | 4.6 GB | 8 GB | 16 GB |
Q6_K | 4.2 GB | 5.3 GB | 5.5 GB | 8 GB | 16 GB |
Q8_0 | 4.8 GB | 5.9 GB | 6.1 GB | 8 GB | 16 GB |
F16 | 8.9 GB | 10.1 GB | 10.3 GB | 12 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Gemma 3n E4B
6.9 billion parameters · up to 32K context · 30 thousand downloads in 30 days
At Q4_K_M, the Gemma 3n E4B file is 4.5 GB, and with 32K tokens of context and runtime overhead it comes to about 5.9 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
A sliding window of 512 tokens on 16 layers keeps the cache small: 0.3 GB at the full 32K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 3.7 GB | 4.8 GB | 5.0 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 4.1 GB | 5.3 GB | 5.5 GB | 8 GB | 16 GB |
Q4_K_M | 4.5 GB | 5.7 GB | 5.9 GB | 8 GB | 16 GB |
Q5_K_M | 5.0 GB | 6.2 GB | 6.4 GB | 8 GB | 16 GB |
Q6_K | 6.3 GB | 7.4 GB | 7.6 GB | 8 GB | 16 GB |
Q8_0 | 7.4 GB | 8.5 GB | 8.7 GB | 12 GB | 16 GB |
F16 | 13.7 GB | 14.9 GB | 15.1 GB | 16 GB | 24 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.