ThinkFacility
  1. Can I run it locally?
  2. Gemma

Gemma VRAM requirements: can I run Gemma 4 locally?

Gemma 4 12B needs 8.9 GB at a 4-bit quant with 32K of context, which a 12 GB card holds. Here's every Gemma in wide use, from the phone-sized E2B to the 31B.

Check your machine

Fits on the card 8.9 GB

Everything stays in video memory, which is the fast case.

  • Weights 7.1 GB
  • Context cache 0.8 GB
  • Overhead 1.0 GB
  • Fast memory
  • With system RAM

What we found

Gemma's cache hardly grows, and that's by design. In Gemma 4 five layers out of every six only look back over a sliding window of 1,024 tokens (512 on the E models), so only one layer in six keeps the whole conversation. Gemma 4 31B at its full 256K context carries 12.0 GB of cache. Gemma 3 27B needs 11.4 GB at half that context.

The E models go further. E4B's last 18 layers reuse keys and values computed by earlier layers, so they add nothing to the cache at all, and its whole 128K context costs 2.2 GB.

Gemma 4 12B is the most downloaded Gemma right now (1.2 million in 30 days) and the sensible pick for a 12 GB card. With 24 GB, Gemma 4 26B A4B fits at 18.6 GB, and since only a few of its experts run for each token it answers faster than the dense 31B.

Google also publishes QAT builds, trained so they lose less at 4-bit. They need the same memory as the regular file at the same quant, so they don't get separate sections here.

Gemma 4 12B

11.9 billion parameters · up to 256K context · 1.2 million downloads in 30 days

Gemma 4 12B needs about 8.9 GB at Q4_K_M with a 32K-token context. The file itself is 7.1 GB, and the cache for 32K tokens adds 0.8 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.

40 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 256K is 2.7 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL4.7 GB6.3 GB6.5 GB8 GB16 GB
UD-Q3_K_XL6.0 GB7.6 GB7.8 GB8 GB16 GB
Q4_K_M7.1 GB8.7 GB8.9 GB12 GB16 GB
Q5_K_M8.4 GB10.0 GB10.2 GB12 GB16 GB
Q6_K9.8 GB11.4 GB11.6 GB12 GB16 GB*
Q8_012.7 GB14.3 GB14.5 GB16 GB24 GB
BF1623.8 GB25.4 GB25.6 GB24 GB32 GB*

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Gemma 4 E4B

7.5 billion parameters · up to 128K context · 594 thousand downloads in 30 days

At Q4_K_M, the Gemma 4 E4B file is 5.0 GB, and with 32K tokens of context and runtime overhead it comes to about 6.6 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.

A sliding window of 512 tokens on 20 layers keeps the cache small: 2.2 GB at the full 128K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL3.8 GB5.0 GB5.4 GB8 GB16 GB
UD-Q3_K_XL4.6 GB5.8 GB6.2 GB8 GB16 GB
Q4_K_M5.0 GB6.2 GB6.6 GB8 GB16 GB
Q5_K_M5.5 GB6.7 GB7.1 GB8 GB16 GB
Q6_K7.1 GB8.3 GB8.7 GB12 GB16 GB
Q8_08.2 GB9.4 GB9.8 GB12 GB16 GB
BF1615.1 GB16.3 GB16.7 GB16 GB24 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 4 26B A4B

25.2 billion parameters · mixture of experts, 128 experts · up to 256K context · 552 thousand downloads in 30 days

Gemma 4 26B A4B needs about 18.6 GB at UD-Q4_K_M with a 32K-token context. The file itself is 16.9 GB, and the cache for 32K tokens adds 0.7 GB. That fits a 24 GB card (RTX 3090, RTX 4090). It routes each token through a few of 128 experts, so it runs faster than its size suggests.

25 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 256K is 3.0 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL10.5 GB12.0 GB12.2 GB12 GB16 GB*
UD-Q3_K_XL12.9 GB14.3 GB14.6 GB16 GB24 GB
UD-Q4_K_M16.9 GB18.4 GB18.6 GB24 GB24 GB*
UD-Q5_K_XL21.2 GB22.7 GB22.9 GB24 GB32 GB
UD-Q6_K_XL23.3 GB24.7 GB25.0 GB24 GB32 GB*
Q8_026.9 GB28.3 GB28.5 GB32 GB32 GB*
BF1650.5 GB51.9 GB52.2 GB96 GB64 GB*

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Gemma 4 31B

30.7 billion parameters · up to 256K context · 371 thousand downloads in 30 days

At Q4_K_M, the Gemma 4 31B file is 18.3 GB, and with 32K tokens of context and runtime overhead it comes to about 22.0 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090). On a Mac, 32 GB of unified memory is enough.

A sliding window of 1,024 tokens on 50 layers keeps the cache small: 12.0 GB at the full 256K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL11.8 GB14.4 GB15.4 GB16 GB24 GB
UD-Q3_K_XL15.4 GB18.0 GB19.0 GB24 GB24 GB*
Q4_K_M18.3 GB21.0 GB22.0 GB24 GB32 GB
Q5_K_M21.7 GB24.3 GB25.3 GB24 GB32 GB*
Q6_K25.2 GB27.8 GB28.8 GB32 GB32 GB*
Q8_032.6 GB35.3 GB36.3 GB48 GB48 GB
BF1661.4 GB64.0 GB65.1 GB96 GB96 GB

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Gemma 3 27B

27.0 billion parameters · up to 128K context · 30 thousand downloads in 30 days

Gemma 3 27B needs about 20.9 GB at Q4_K_M with a 32K-token context. The file itself is 16.5 GB, and the cache for 32K tokens adds 3.3 GB. That fits a 24 GB card (RTX 3090, RTX 4090).

52 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 11.4 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_S6.5 GB8.9 GB10.9 GB12 GB16 GB
UD-Q2_K_XL10.7 GB13.0 GB15.1 GB16 GB24 GB
UD-Q3_K_XL13.7 GB16.1 GB18.1 GB24 GB24 GB*
Q4_K_M16.5 GB18.9 GB20.9 GB24 GB24 GB*
Q5_K_M19.3 GB21.6 GB23.6 GB24 GB32 GB*
Q6_K22.2 GB24.5 GB26.5 GB32 GB32 GB*
Q8_028.7 GB31.1 GB33.1 GB32 GB36 GB*
BF1654.0 GB56.4 GB58.4 GB96 GB64 GB*

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Gemma 4 E2B

4.6 billion parameters · up to 128K context · 452 thousand downloads in 30 days

At Q4_K_M, the Gemma 4 E2B file is 3.1 GB, and with 32K tokens of context and runtime overhead it comes to about 4.4 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.

A sliding window of 512 tokens on 12 layers keeps the cache small: 0.8 GB at the full 128K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL2.4 GB3.5 GB3.7 GB8 GB16 GB
UD-Q3_K_XL2.9 GB4.0 GB4.2 GB8 GB16 GB
Q4_K_M3.1 GB4.2 GB4.4 GB8 GB16 GB
Q5_K_M3.4 GB4.5 GB4.6 GB8 GB16 GB
Q6_K4.5 GB5.6 GB5.8 GB8 GB16 GB
Q8_05.0 GB6.1 GB6.3 GB8 GB16 GB
BF169.3 GB10.4 GB10.6 GB12 GB16 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 3 4B

3.9 billion parameters · up to 128K context · 84 thousand downloads in 30 days

Gemma 3 4B needs about 4.4 GB at Q4_K_M with a 32K-token context. The file itself is 2.5 GB, and the cache for 32K tokens adds 0.9 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.

29 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 2.9 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_S1.2 GB2.6 GB3.1 GB8 GB16 GB
UD-Q2_K_XL1.8 GB3.2 GB3.7 GB8 GB16 GB
UD-Q3_K_XL2.2 GB3.5 GB4.0 GB8 GB16 GB
Q4_K_M2.5 GB3.9 GB4.4 GB8 GB16 GB
Q5_K_M2.8 GB4.2 GB4.7 GB8 GB16 GB
Q6_K3.2 GB4.6 GB5.1 GB8 GB16 GB
Q8_04.1 GB5.5 GB6.0 GB8 GB16 GB
BF167.8 GB9.2 GB9.7 GB12 GB16 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 3 270M

268 million parameters · up to 32K context · 63 thousand downloads in 30 days

At Q4_K_M, the Gemma 3 270M file is 0.3 GB, and with 32K tokens of context and runtime overhead it comes to about 1.4 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.

A sliding window of 512 tokens on 15 layers keeps the cache small: 0.1 GB at the full 32K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL0.2 GB1.3 GB1.4 GB8 GB16 GB
UD-Q3_K_XL0.2 GB1.3 GB1.4 GB8 GB16 GB
Q4_K_M0.3 GB1.3 GB1.4 GB8 GB16 GB
Q5_K_M0.3 GB1.3 GB1.4 GB8 GB16 GB
Q6_K0.3 GB1.4 GB1.4 GB8 GB16 GB
Q8_00.3 GB1.4 GB1.4 GB8 GB16 GB
F160.5 GB1.6 GB1.7 GB8 GB16 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 3 12B

11.8 billion parameters · up to 128K context · 62 thousand downloads in 30 days

Gemma 3 12B needs about 11.0 GB at Q4_K_M with a 32K-token context. The file itself is 7.3 GB, and the cache for 32K tokens adds 2.7 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.

40 of its layers only look back 1,024 tokens, so they stop growing early, and the whole cache at 128K is 9.1 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_S3.1 GB5.1 GB6.7 GB8 GB16 GB
UD-Q2_K_XL4.9 GB6.9 GB8.5 GB8 GB16 GB
UD-Q3_K_XL6.1 GB8.2 GB9.8 GB12 GB16 GB
Q4_K_M7.3 GB9.4 GB11.0 GB12 GB16 GB
Q5_K_M8.4 GB10.5 GB12.1 GB12 GB16 GB*
Q6_K9.7 GB11.7 GB13.3 GB16 GB24 GB
Q8_012.5 GB14.6 GB16.2 GB16 GB24 GB
BF1623.5 GB25.6 GB27.2 GB32 GB32 GB*

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Gemma 3 1B

1000 million parameters · up to 32K context · 42 thousand downloads in 30 days

At Q4_K_M, the Gemma 3 1B file is 0.8 GB, and with 32K tokens of context and runtime overhead it comes to about 2.0 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.

A sliding window of 512 tokens on 22 layers keeps the cache small: 0.2 GB at the full 32K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_S0.6 GB1.7 GB1.8 GB8 GB16 GB
UD-Q2_K_XL0.7 GB1.8 GB1.9 GB8 GB16 GB
UD-Q3_K_XL0.7 GB1.8 GB1.9 GB8 GB16 GB
Q4_K_M0.8 GB1.9 GB2.0 GB8 GB16 GB
Q5_K_M0.9 GB1.9 GB2.0 GB8 GB16 GB
Q6_K1.0 GB2.1 GB2.2 GB8 GB16 GB
Q8_01.1 GB2.2 GB2.3 GB8 GB16 GB
BF162.0 GB3.1 GB3.2 GB8 GB16 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 3n E2B

4.5 billion parameters · up to 32K context · 39 thousand downloads in 30 days

Gemma 3n E2B needs about 4.4 GB at Q4_K_M with a 32K-token context. The file itself is 3.0 GB, and the cache for 32K tokens adds 0.3 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.

16 of its layers only look back 512 tokens, so they stop growing early, and the whole cache at 32K is 0.3 GB.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL2.6 GB3.8 GB4.0 GB8 GB16 GB
UD-Q3_K_XL2.8 GB4.0 GB4.2 GB8 GB16 GB
Q4_K_M3.0 GB4.2 GB4.4 GB8 GB16 GB
Q5_K_M3.3 GB4.4 GB4.6 GB8 GB16 GB
Q6_K4.2 GB5.3 GB5.5 GB8 GB16 GB
Q8_04.8 GB5.9 GB6.1 GB8 GB16 GB
F168.9 GB10.1 GB10.3 GB12 GB16 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Gemma 3n E4B

6.9 billion parameters · up to 32K context · 30 thousand downloads in 30 days

At Q4_K_M, the Gemma 3n E4B file is 4.5 GB, and with 32K tokens of context and runtime overhead it comes to about 5.9 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.

A sliding window of 512 tokens on 16 layers keeps the cache small: 0.3 GB at the full 32K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL3.7 GB4.8 GB5.0 GB8 GB16 GB
UD-Q3_K_XL4.1 GB5.3 GB5.5 GB8 GB16 GB
Q4_K_M4.5 GB5.7 GB5.9 GB8 GB16 GB
Q5_K_M5.0 GB6.2 GB6.4 GB8 GB16 GB
Q6_K6.3 GB7.4 GB7.6 GB8 GB16 GB
Q8_07.4 GB8.5 GB8.7 GB12 GB16 GB
F1613.7 GB14.9 GB15.1 GB16 GB24 GB

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

How the numbers are worked out

Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:

cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens

Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.

A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.

Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.