Qwen VRAM requirements: can I run Qwen3.8 27B locally?
Qwen3.8 27B at a 4-bit quant with a 32K-token context comes to 19.8 GB, so it fits a 24 GB card. Below, every Qwen size people are downloading, each quant, and what your own machine can hold.
Check your machine
Fits on the card 22.6 GB
Everything stays in video memory, which is the fast case.
- Weights 18.6 GB
- Context cache 3.2 GB
- Overhead 0.8 GB
- Fast memory
- With system RAM
What we found
The number I didn't expect is the cache. Qwen3.6 35B A3B is the bigger model, and yet at its full 256K context it needs 28.6 GB, where the older Qwen3 Coder 30B A3B needs 45.1 GB for the same 256K. Since Qwen3.5, three layers in every four use linear attention, which keeps a small state of fixed size instead of a record of every token it has read. Long context got cheap.
On a 24 GB card the one to start with is Qwen3.8 27B. At UD-Q4_K_M it's 19.8 GB with a 32K context, and doubling the context to 64K only takes it to 21.9 GB. It's also what people are fetching: 6.9 million downloads of Unsloth's files in 30 days, beaten only by the Coder model from July 2025 (still at 11.6 million, which surprised me).
The A3B models are the ones for a smaller card. They hold 35 billion parameters but route each token through about 3 billion of them, so llama.cpp can leave the experts in system RAM and run the rest on a 12 GB GPU at a speed people live with. The dense 27B can't do that trick, because every token touches every weight.
Qwen3.8 Flash Next is another class of machine: 177 billion parameters and 111 GB at UD-Q4_K_XL. That's a DGX Spark, or a 128 GB Mac with its GPU memory limit raised. Our Qwen guide has which of these are open and which are API only.
Qwen3 Coder 30B A3B
30.5 billion parameters · mixture of experts, 8 of 128 active · up to 256K context · 11.6 million downloads in 30 days
Qwen3 Coder 30B A3B needs about 22.6 GB at Q4_K_M with a 32K-token context. The file itself is 18.6 GB, and the cache for 32K tokens adds 3.2 GB. That fits a 24 GB card (RTX 3090, RTX 4090). On a Mac, 32 GB of unified memory is enough. It routes each token through 8 of 128 experts, so it runs faster than its size suggests.
Every token of context adds 96 KB of cache, so the full 256K would put 25.8 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-TQ1_0 | 8.0 GB | 9.6 GB | 12.0 GB | 12 GB | 16 GB* |
UD-Q2_K_XL | 11.8 GB | 13.4 GB | 15.8 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 13.8 GB | 15.4 GB | 17.8 GB | 24 GB | 24 GB* |
Q4_K_M | 18.6 GB | 20.2 GB | 22.6 GB | 24 GB | 32 GB |
Q5_K_M | 21.7 GB | 23.3 GB | 25.8 GB | 24 GB | 32 GB* |
Q6_K | 25.1 GB | 26.7 GB | 29.1 GB | 32 GB | 32 GB* |
Q8_0 | 32.5 GB | 34.1 GB | 36.5 GB | 48 GB | 48 GB |
BF16 | 61.1 GB | 62.7 GB | 65.1 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.8 27B
27.3 billion parameters · up to 256K context · 6.9 million downloads in 30 days
At UD-Q4_K_M, the Qwen3.8 27B file is 16.5 GB, and with 32K tokens of context and runtime overhead it comes to about 19.8 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090).
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 64 KB a token and reaches 17.3 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 6.2 GB | 7.9 GB | 9.5 GB | 12 GB | 16 GB |
UD-Q2_K_XL | 9.8 GB | 11.5 GB | 13.1 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 13.1 GB | 14.8 GB | 16.5 GB | 16 GB | 24 GB |
UD-Q4_K_M | 16.5 GB | 18.2 GB | 19.8 GB | 24 GB | 24 GB* |
UD-Q5_K_XL | 20.9 GB | 22.6 GB | 24.2 GB | 24 GB | 32 GB* |
UD-Q6_K_XL | 25.3 GB | 27.0 GB | 28.6 GB | 32 GB | 32 GB* |
Q8_0 | 29.0 GB | 30.7 GB | 32.4 GB | 32 GB | 36 GB* |
BF16 | 54.7 GB | 56.4 GB | 58.0 GB | 96 GB | 64 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.8 Flash Next
177 billion parameters · mixture of experts, 10 of 512 active · up to 256K context · 1.8 million downloads in 30 days
Qwen3.8 Flash Next needs about 113 GB at UD-Q4_K_XL with a 32K-token context. The file itself is 111 GB, and the cache for 32K tokens adds 0.9 GB. No single consumer card holds it. A Mac with 128 GB of unified memory does once you raise the GPU memory limit, and a PC can split it between a GPU and system RAM. It's a mixture-of-experts model (10 of 512 experts per token), which is why people run it with the experts parked in system RAM and the rest on a 24 GB card.
Only 12 of its 48 layers keep a growing cache (the rest carry a fixed-size state), so the full 256K context costs 6.6 GB of cache.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 72.5 GB | 73.9 GB | 74.5 GB | 96 GB | 96 GB |
UD-Q2_K_XL | 78.9 GB | 80.2 GB | 80.8 GB | 96 GB | 96 GB* |
UD-Q3_K_XL | 90.0 GB | 91.3 GB | 91.9 GB | 96 GB | 96 GB* |
UD-Q4_K_XL | 111 GB | 113 GB | 113 GB | none | 128 GB* |
UD-Q5_K_XL | 158 GB | 160 GB | 160 GB | none | 256 GB |
UD-Q6_K_XL | 169 GB | 170 GB | 171 GB | none | 256 GB |
Q8_0 | 188 GB | 190 GB | 190 GB | none | 256 GB |
BF16 | 354 GB | 355 GB | 356 GB | none | 512 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.6 35B A3B
34.7 billion parameters · mixture of experts, 8 of 256 active · up to 256K context · 1.3 million downloads in 30 days
At UD-Q4_K_M, the Qwen3.6 35B A3B file is 22.1 GB, and with 32K tokens of context and runtime overhead it comes to about 23.9 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090). It routes each token through 8 of 256 experts, so it runs faster than its size suggests.
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 20 KB a token and reaches 5.4 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_M | 10.0 GB | 11.3 GB | 11.8 GB | 12 GB | 16 GB* |
UD-Q2_K_XL | 12.3 GB | 13.5 GB | 14.0 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 16.8 GB | 18.1 GB | 18.6 GB | 24 GB | 24 GB* |
UD-Q4_K_M | 22.1 GB | 23.4 GB | 23.9 GB | 24 GB | 32 GB* |
UD-Q5_K_XL | 26.6 GB | 27.8 GB | 28.3 GB | 32 GB | 32 GB* |
UD-Q6_K_XL | 31.8 GB | 33.1 GB | 33.6 GB | 32 GB | 36 GB* |
Q8_0 | 36.9 GB | 38.1 GB | 38.6 GB | 48 GB | 48 GB |
BF16 | 69.4 GB | 70.6 GB | 71.1 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.5 9B
9.0 billion parameters · up to 256K context · 1.4 million downloads in 30 days
Qwen3.5 9B needs about 7.8 GB at Q4_K_M with a 32K-token context. The file itself is 5.7 GB, and the cache for 32K tokens adds 1.1 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
Only 8 of its 32 layers keep a growing cache (the rest carry a fixed-size state), so the full 256K context costs 8.6 GB of cache.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 4.1 GB | 5.5 GB | 6.3 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 5.1 GB | 6.4 GB | 7.2 GB | 8 GB | 16 GB |
Q4_K_M | 5.7 GB | 7.0 GB | 7.8 GB | 8 GB | 16 GB |
Q5_K_M | 6.6 GB | 7.9 GB | 8.7 GB | 12 GB | 16 GB |
Q6_K | 7.5 GB | 8.8 GB | 9.6 GB | 12 GB | 16 GB |
Q8_0 | 9.5 GB | 10.9 GB | 11.7 GB | 12 GB | 16 GB* |
BF16 | 17.9 GB | 19.3 GB | 20.1 GB | 24 GB | 24 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.6 27B
26.9 billion parameters · up to 256K context · 981 thousand downloads in 30 days
At Q4_K_M, the Qwen3.6 27B file is 16.8 GB, and with 32K tokens of context and runtime overhead it comes to about 20.1 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090).
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 64 KB a token and reaches 17.3 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 11.8 GB | 13.6 GB | 15.2 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 14.5 GB | 16.2 GB | 17.8 GB | 24 GB | 24 GB* |
Q4_K_M | 16.8 GB | 18.5 GB | 20.1 GB | 24 GB | 24 GB* |
Q5_K_M | 19.5 GB | 21.2 GB | 22.8 GB | 24 GB | 32 GB |
Q6_K | 22.5 GB | 24.2 GB | 25.8 GB | 32 GB | 32 GB* |
Q8_0 | 28.6 GB | 30.3 GB | 31.9 GB | 32 GB | 36 GB* |
BF16 | 53.8 GB | 55.5 GB | 57.1 GB | 96 GB | 64 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.5 4B
4.2 billion parameters · up to 256K context · 922 thousand downloads in 30 days
Qwen3.5 4B needs about 4.9 GB at Q4_K_M with a 32K-token context. The file itself is 2.7 GB, and the cache for 32K tokens adds 1.1 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
Only 8 of its 32 layers keep a growing cache (the rest carry a fixed-size state), so the full 256K context costs 8.6 GB of cache.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 1.9 GB | 3.3 GB | 4.1 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.4 GB | 3.8 GB | 4.6 GB | 8 GB | 16 GB |
Q4_K_M | 2.7 GB | 4.1 GB | 4.9 GB | 8 GB | 16 GB |
Q5_K_M | 3.1 GB | 4.5 GB | 5.3 GB | 8 GB | 16 GB |
Q6_K | 3.5 GB | 4.9 GB | 5.7 GB | 8 GB | 16 GB |
Q8_0 | 4.5 GB | 5.8 GB | 6.6 GB | 8 GB | 16 GB |
BF16 | 8.4 GB | 9.8 GB | 10.6 GB | 12 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Qwen3 30B A3B 2507
30.5 billion parameters · mixture of experts, 8 of 128 active · up to 256K context · 521 thousand downloads in 30 days
At Q4_K_M, the Qwen3 30B A3B 2507 file is 18.6 GB, and with 32K tokens of context and runtime overhead it comes to about 22.6 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090). On a Mac, 32 GB of unified memory is enough. It routes each token through 8 of 128 experts, so it runs faster than its size suggests.
The cache is the part that grows: 96 KB a token, or 25.8 GB if you use all 256K tokens it supports.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-TQ1_0 | 8.1 GB | 9.7 GB | 12.1 GB | 12 GB | 16 GB* |
UD-Q2_K_XL | 11.8 GB | 13.4 GB | 15.8 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 13.8 GB | 15.4 GB | 17.9 GB | 24 GB | 24 GB* |
Q4_K_M | 18.6 GB | 20.2 GB | 22.6 GB | 24 GB | 32 GB |
Q5_K_M | 21.7 GB | 23.3 GB | 25.8 GB | 24 GB | 32 GB* |
Q6_K | 25.1 GB | 26.7 GB | 29.1 GB | 32 GB | 32 GB* |
Q8_0 | 32.5 GB | 34.1 GB | 36.5 GB | 48 GB | 48 GB |
BF16 | 61.1 GB | 62.7 GB | 65.1 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3 4B
4.0 billion parameters · up to 40K context · 462 thousand downloads in 30 days
Qwen3 4B needs about 8.1 GB at Q4_K_M with a 32K-token context. The file itself is 2.5 GB, and the cache for 32K tokens adds 4.8 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
Every token of context adds 144 KB of cache, so the full 40K would put 6.0 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 1.1 GB | 3.1 GB | 6.7 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 1.7 GB | 3.7 GB | 7.3 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.1 GB | 4.2 GB | 7.8 GB | 8 GB | 16 GB |
Q4_K_M | 2.5 GB | 4.5 GB | 8.1 GB | 8 GB | 16 GB |
Q5_K_M | 2.9 GB | 4.9 GB | 8.5 GB | 8 GB | 16 GB |
Q6_K | 3.3 GB | 5.3 GB | 8.9 GB | 12 GB | 16 GB |
Q8_0 | 4.3 GB | 6.3 GB | 9.9 GB | 12 GB | 16 GB |
BF16 | 8.1 GB | 10.1 GB | 13.7 GB | 16 GB | 24 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Qwen AgentWorld 35B A3B
34.7 billion parameters · mixture of experts, 8 of 256 active · up to 256K context · 382 thousand downloads in 30 days
At UD-Q4_K_M, the Qwen AgentWorld 35B A3B file is 22.1 GB, and with 32K tokens of context and runtime overhead it comes to about 23.9 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090). It routes each token through 8 of 256 experts, so it runs faster than its size suggests.
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 20 KB a token and reaches 5.4 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 12.3 GB | 13.5 GB | 14.0 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 16.8 GB | 18.0 GB | 18.5 GB | 24 GB | 24 GB* |
UD-Q4_K_M | 22.1 GB | 23.4 GB | 23.9 GB | 24 GB | 32 GB* |
UD-Q5_K_XL | 26.5 GB | 27.8 GB | 28.3 GB | 32 GB | 32 GB* |
UD-Q6_K_XL | 31.8 GB | 33.1 GB | 33.6 GB | 32 GB | 36 GB* |
Q8_0 | 36.9 GB | 38.1 GB | 38.6 GB | 48 GB | 48 GB |
BF16 | 69.4 GB | 70.6 GB | 71.1 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3.5 2B
1.9 billion parameters · up to 256K context · 319 thousand downloads in 30 days
Qwen3.5 2B needs about 2.7 GB at Q4_K_M with a 32K-token context. The file itself is 1.3 GB, and the cache for 32K tokens adds 0.4 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
Only 6 of its 24 layers keep a growing cache (the rest carry a fixed-size state), so the full 256K context costs 3.2 GB of cache.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 1.0 GB | 2.1 GB | 2.4 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 1.2 GB | 2.3 GB | 2.6 GB | 8 GB | 16 GB |
Q4_K_M | 1.3 GB | 2.4 GB | 2.7 GB | 8 GB | 16 GB |
Q5_K_M | 1.4 GB | 2.6 GB | 2.9 GB | 8 GB | 16 GB |
Q6_K | 1.6 GB | 2.7 GB | 3.0 GB | 8 GB | 16 GB |
Q8_0 | 2.0 GB | 3.1 GB | 3.4 GB | 8 GB | 16 GB |
BF16 | 3.8 GB | 4.9 GB | 5.2 GB | 8 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Qwen3.5 27B
26.9 billion parameters · up to 256K context · 314 thousand downloads in 30 days
At Q4_K_M, the Qwen3.5 27B file is 16.7 GB, and with 32K tokens of context and runtime overhead it comes to about 20.1 GB. The smallest card that holds it whole is a 24 GB card (RTX 3090, RTX 4090).
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 64 KB a token and reaches 17.3 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 11.2 GB | 12.9 GB | 14.5 GB | 16 GB | 24 GB |
UD-Q3_K_XL | 14.4 GB | 16.1 GB | 17.8 GB | 24 GB | 24 GB* |
Q4_K_M | 16.7 GB | 18.4 GB | 20.1 GB | 24 GB | 24 GB* |
Q5_K_M | 19.6 GB | 21.3 GB | 22.9 GB | 24 GB | 32 GB* |
Q6_K | 22.5 GB | 24.2 GB | 25.8 GB | 24 GB | 32 GB* |
Q8_0 | 28.6 GB | 30.3 GB | 31.9 GB | 32 GB | 36 GB* |
BF16 | 53.8 GB | 55.5 GB | 57.1 GB | 96 GB | 64 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Qwen3 VL 4B
4.0 billion parameters · up to 256K context · 314 thousand downloads in 30 days
Qwen3 VL 4B needs about 8.1 GB at Q4_K_M with a 32K-token context. The file itself is 2.5 GB, and the cache for 32K tokens adds 4.8 GB. That fits a 8 GB card (RTX 4060, RTX 3070). On a Mac, 16 GB of unified memory is enough.
Every token of context adds 144 KB of cache, so the full 256K would put 38.7 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 1.1 GB | 3.1 GB | 6.7 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 1.7 GB | 3.7 GB | 7.3 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 2.1 GB | 4.1 GB | 7.8 GB | 8 GB | 16 GB |
Q4_K_M | 2.5 GB | 4.5 GB | 8.1 GB | 8 GB | 16 GB |
Q5_K_M | 2.9 GB | 4.9 GB | 8.5 GB | 8 GB | 16 GB |
Q6_K | 3.3 GB | 5.3 GB | 8.9 GB | 12 GB | 16 GB |
Q8_0 | 4.3 GB | 6.3 GB | 9.9 GB | 12 GB | 16 GB |
BF16 | 8.1 GB | 10.1 GB | 13.7 GB | 16 GB | 24 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Qwen3.5 122B A10B
122 billion parameters · mixture of experts, 8 of 256 active · up to 256K context · 261 thousand downloads in 30 days
At Q4_K_M, the Qwen3.5 122B A10B file is 76.5 GB, and with 32K tokens of context and runtime overhead it comes to about 78.5 GB. The smallest card that holds it whole is a RTX PRO 6000 (96 GB). It's a mixture-of-experts model (8 of 256 experts per token), which is why people run it with the experts parked in system RAM and the rest on a 24 GB card.
Most of its layers use linear attention, which keeps a fixed state instead of a cache, so the cache grows by about 24 KB a token and reaches 6.6 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_M | 34.2 GB | 35.6 GB | 36.2 GB | 48 GB | 48 GB |
UD-Q2_K_XL | 41.8 GB | 43.2 GB | 43.8 GB | 48 GB | 48 GB* |
UD-Q3_K_XL | 57.0 GB | 58.3 GB | 58.9 GB | 96 GB | 64 GB* |
Q4_K_M | 76.5 GB | 77.9 GB | 78.5 GB | 96 GB | 96 GB* |
Q5_K_M | 91.5 GB | 92.9 GB | 93.5 GB | 96 GB | 96 GB* |
Q6_K | 101 GB | 102 GB | 103 GB | 96 GB | 128 GB |
Q8_0 | 130 GB | 131 GB | 132 GB | none | 128 GB* |
BF16 | 244 GB | 246 GB | 246 GB | none | 256 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.