- Can I run it locally?
- DeepSeek
DeepSeek VRAM requirements: can I run DeepSeek V4 Flash locally?
DeepSeek V4 Flash is 155 GB at UD-Q4_K_XL, and its smallest file is UD-IQ1_S (82.5 GB). Here's which machines hold each DeepSeek model, down to the R1 distills that fit a gaming PC.
Check your machine
Won't fit 156 GB
Weights, cache and overhead add up to more than the card and the RAM together.
- Weights 155 GB
- Context cache not computed
- Overhead 0.8 GB
- Fast memory
- With system RAM
What we found
DeepSeek's current models are big. DeepSeek V4 Flash 0731 has 284 billion parameters. The files that fit a 128 GB Mac or a DGX Spark are the 1-bit and 2-bit ones, starting at UD-IQ1_S (82.5 GB). At 4-bit you're looking at a 256 GB Mac.
Those V4 totals leave the context cache out, and I'd rather say so than guess. V4 compresses its attention cache in a way its config describes with per-layer compression ratios, and we haven't matched that to how llama.cpp stores it. DeepSeek's own figure points small: its V4.1 Flash announcement puts that model at 890 bytes a token and calls it about a quarter of V4 Flash.
DeepSeek V4 Pro 0813 is past any single desktop. Its smallest file is UD-Q4_K_XL (850 GB), bigger than a 512 GB Mac.
Most of the smaller DeepSeek downloads are the R1 distills from January 2025: Qwen and Llama models fine-tuned on R1's reasoning, which run exactly like the models they're built on. DeepSeek R1 Distill Qwen 14B at 16.2 GB is the one for a 16 GB card.
DeepSeek V4 Flash 0731
284 billion parameters · mixture of experts, 6 of 256 active · up to 1M context · 198 thousand downloads in 30 days
DeepSeek V4 Flash 0731 needs about 156 GB at UD-Q4_K_XL before its context cache: the file is 155 GB and llama.cpp wants roughly 0.8 GB more to run it. No single consumer card holds it. A Mac with 256 GB of unified memory does, and a PC can split it between a GPU and system RAM. It's a mixture-of-experts model (6 of 256 experts per token), which is why people run it with the experts parked in system RAM and the rest on a 24 GB card.
We don't compute the context cache for the deepseek_v4 architecture yet, because its config describes an attention scheme we haven't matched to llama.cpp's cache, so the totals here are weights and overhead only.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 82.5 GB | 83.3 GB | 83.3 GB | 96 GB | 96 GB* |
UD-Q2_K_XL | 96.8 GB | 97.6 GB | 97.6 GB | 96 GB | 96 GB* |
UD-Q3_K_XL | 128 GB | 129 GB | 129 GB | none | 128 GB* |
UD-Q4_K_XL | 155 GB | 156 GB | 156 GB | none | 256 GB |
UD-Q8_K_XL | 162 GB | 163 GB | 163 GB | none | 256 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek V4 Pro 0813
1.57 trillion parameters · mixture of experts, 6 of 384 active · up to 1M context · 138 thousand downloads in 30 days
DeepSeek V4 Pro 0813 needs about 850 GB at UD-Q4_K_XL before its context cache: the file is 850 GB and llama.cpp wants roughly 0.8 GB more to run it. That's past a 512 GB Mac. Even the smallest file, UD-Q4_K_XL, is 850 GB, so this one is for servers with several GPUs, or a workstation with a lot of system RAM and patience. It routes each token through 6 of 384 experts, so it runs faster than its size suggests.
We don't compute the context cache for the deepseek_v4 architecture yet, because its config describes an attention scheme we haven't matched to llama.cpp's cache, so the totals here are weights and overhead only.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q4_K_XL | 850 GB | 850 GB | 850 GB | none | none |
UD-Q8_K_XL | 873 GB | 874 GB | 874 GB | none | none |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
DeepSeek R1 Distill Qwen 14B
14.8 billion parameters · up to 128K context · 48 thousand downloads in 30 days
DeepSeek R1 Distill Qwen 14B needs about 16.2 GB at Q4_K_M with a 32K-token context. The file itself is 9.0 GB, and the cache for 32K tokens adds 6.4 GB. That fits a 16 GB card (RTX 4060 Ti 16 GB, 5070 Ti, 5080). On a Mac, 24 GB of unified memory is enough.
Every token of context adds 192 KB of cache, so the full 128K would put 25.8 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
Q2_K | 5.8 GB | 8.2 GB | 13.0 GB | 16 GB | 24 GB |
Q3_K_M | 7.3 GB | 9.8 GB | 14.6 GB | 16 GB | 24 GB |
Q4_K_M | 9.0 GB | 11.4 GB | 16.2 GB | 16 GB | 24 GB |
Q5_K_M | 10.5 GB | 12.9 GB | 17.8 GB | 24 GB | 24 GB* |
Q6_K | 12.1 GB | 14.5 GB | 19.4 GB | 24 GB | 24 GB* |
Q8_0 | 15.7 GB | 18.1 GB | 23.0 GB | 24 GB | 32 GB* |
F16 | 29.5 GB | 32.0 GB | 36.8 GB | 48 GB | 48 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1
671 billion parameters · mixture of experts, 8 of 256 active · up to 160K context · 29 thousand downloads in 30 days
At Q4_K_M, the DeepSeek R1 file is 404 GB, and with 32K tokens of context and runtime overhead it comes to about 407 GB. No single consumer card holds it. A Mac with 512 GB of unified memory does, and a PC can split it between a GPU and system RAM. It routes each token through 8 of 256 experts, so it runs faster than its size suggests.
It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 11.5 GB at 160K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 140 GB | 142 GB | 143 GB | none | 256 GB |
UD-Q2_K_XL | 227 GB | 228 GB | 230 GB | none | 256 GB* |
Q3_K_M | 319 GB | 321 GB | 322 GB | none | 512 GB |
Q4_K_M | 404 GB | 406 GB | 407 GB | none | 512 GB |
Q5_K_M | 475 GB | 477 GB | 478 GB | none | 512 GB* |
Q6_K | 551 GB | 552 GB | 554 GB | none | none |
Q8_0 | 713 GB | 715 GB | 716 GB | none | none |
BF16 | 1,342 GB | 1,344 GB | 1,345 GB | none | none |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1 Distill Llama 70B
70.6 billion parameters · up to 128K context · 138 thousand downloads in 30 days
DeepSeek R1 Distill Llama 70B needs about 54.0 GB at Q4_K_M with a 32K-token context. The file itself is 42.5 GB, and the cache for 32K tokens adds 10.7 GB. That fits a RTX PRO 6000 (96 GB).
Every token of context adds 320 KB of cache, so the full 128K would put 42.9 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 15.9 GB | 19.3 GB | 27.4 GB | 32 GB | 32 GB* |
UD-Q2_K_XL | 27.0 GB | 30.4 GB | 38.5 GB | 48 GB | 48 GB |
UD-Q3_K_XL | 34.9 GB | 38.3 GB | 46.4 GB | 48 GB | 48 GB* |
Q4_K_M | 42.5 GB | 46.0 GB | 54.0 GB | 96 GB | 64 GB* |
Q5_K_M | 49.9 GB | 53.4 GB | 61.4 GB | 96 GB | 64 GB* |
Q6_K | 116 GB | 119 GB | 127 GB | none | 128 GB* |
Q8_0 | 75.0 GB | 78.4 GB | 86.5 GB | 96 GB | 96 GB* |
BF16 | 141 GB | 145 GB | 153 GB | none | 256 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1 Distill Qwen 1.5B
1.8 billion parameters · up to 128K context · 58 thousand downloads in 30 days
At Q4_K_M, the DeepSeek R1 Distill Qwen 1.5B file is 1.1 GB, and with 32K tokens of context and runtime overhead it comes to about 2.9 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
The cache is the part that grows: 28 KB a token, or 3.8 GB if you use all 128K tokens it supports.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 0.7 GB | 1.7 GB | 2.5 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 0.9 GB | 1.9 GB | 2.6 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 1.0 GB | 2.0 GB | 2.7 GB | 8 GB | 16 GB |
Q4_K_M | 1.1 GB | 2.2 GB | 2.9 GB | 8 GB | 16 GB |
Q5_K_M | 1.3 GB | 2.3 GB | 3.0 GB | 8 GB | 16 GB |
Q6_K | 1.5 GB | 2.5 GB | 3.2 GB | 8 GB | 16 GB |
Q8_0 | 1.9 GB | 2.9 GB | 3.6 GB | 8 GB | 16 GB |
BF16 | 3.6 GB | 4.6 GB | 5.3 GB | 8 GB | 16 GB |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
DeepSeek R1 0528 Qwen3 8B
8.2 billion parameters · up to 128K context · 57 thousand downloads in 30 days
DeepSeek R1 0528 Qwen3 8B needs about 10.7 GB at Q4_K_M with a 32K-token context. The file itself is 5.0 GB, and the cache for 32K tokens adds 4.8 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.
Every token of context adds 144 KB of cache, so the full 128K would put 19.3 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 2.3 GB | 4.3 GB | 7.9 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 3.5 GB | 5.5 GB | 9.1 GB | 12 GB | 16 GB |
UD-Q3_K_XL | 4.3 GB | 6.3 GB | 10.0 GB | 12 GB | 16 GB |
Q4_K_M | 5.0 GB | 7.0 GB | 10.7 GB | 12 GB | 16 GB |
Q5_K_M | 5.9 GB | 7.9 GB | 11.5 GB | 12 GB | 16 GB* |
Q6_K | 6.7 GB | 8.7 GB | 12.4 GB | 12 GB | 16 GB* |
Q8_0 | 8.7 GB | 10.7 GB | 14.4 GB | 16 GB | 24 GB |
BF16 | 16.4 GB | 18.4 GB | 22.0 GB | 24 GB | 32 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1 Distill Qwen 32B
32.8 billion parameters · up to 128K context · 56 thousand downloads in 30 days
At Q4_K_M, the DeepSeek R1 Distill Qwen 32B file is 19.9 GB, and with 32K tokens of context and runtime overhead it comes to about 29.3 GB. The smallest card that holds it whole is a RTX 5090 (32 GB).
The cache is the part that grows: 256 KB a token, or 34.4 GB if you use all 128K tokens it supports.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
Q2_K | 12.3 GB | 15.3 GB | 21.7 GB | 24 GB | 32 GB |
Q3_K_M | 15.9 GB | 18.9 GB | 25.3 GB | 24 GB | 32 GB* |
Q4_K_M | 19.9 GB | 22.8 GB | 29.3 GB | 32 GB | 32 GB* |
Q5_K_M | 23.3 GB | 26.2 GB | 32.7 GB | 32 GB | 36 GB* |
Q6_K | 26.9 GB | 29.8 GB | 36.3 GB | 48 GB | 48 GB |
Q8_0 | 34.8 GB | 37.8 GB | 44.2 GB | 48 GB | 48 GB* |
F16 | 65.5 GB | 68.5 GB | 74.9 GB | 96 GB | 96 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1 Distill Llama 8B
8.0 billion parameters · up to 128K context · 43 thousand downloads in 30 days
DeepSeek R1 Distill Llama 8B needs about 10.0 GB at Q4_K_M with a 32K-token context. The file itself is 4.9 GB, and the cache for 32K tokens adds 4.3 GB. That fits a RTX 3060 12 GB. On a Mac, 16 GB of unified memory is enough.
Every token of context adds 128 KB of cache, so the full 128K would put 17.2 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 2.2 GB | 4.0 GB | 7.2 GB | 8 GB | 16 GB |
UD-Q2_K_XL | 3.4 GB | 5.2 GB | 8.4 GB | 8 GB | 16 GB |
UD-Q3_K_XL | 4.2 GB | 6.0 GB | 9.3 GB | 12 GB | 16 GB |
Q4_K_M | 4.9 GB | 6.8 GB | 10.0 GB | 12 GB | 16 GB |
Q5_K_M | 5.7 GB | 7.6 GB | 10.8 GB | 12 GB | 16 GB |
Q6_K | 6.6 GB | 8.4 GB | 11.7 GB | 12 GB | 16 GB* |
Q8_0 | 8.5 GB | 10.4 GB | 13.6 GB | 16 GB | 24 GB |
BF16 | 16.1 GB | 17.9 GB | 21.1 GB | 24 GB | 24 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
DeepSeek R1 Distill Qwen 7B
7.6 billion parameters · up to 128K context · 29 thousand downloads in 30 days
At Q4_K_M, the DeepSeek R1 Distill Qwen 7B file is 4.7 GB, and with 32K tokens of context and runtime overhead it comes to about 7.4 GB. Even a 8 GB card (RTX 4060, RTX 3070) holds that. On a Mac, 16 GB of unified memory is enough.
The cache is the part that grows: 56 KB a token, or 7.5 GB if you use all 128K tokens it supports.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
Q2_K | 3.0 GB | 4.3 GB | 5.7 GB | 8 GB | 16 GB |
Q3_K_M | 3.8 GB | 5.1 GB | 6.5 GB | 8 GB | 16 GB |
Q4_K_M | 4.7 GB | 6.0 GB | 7.4 GB | 8 GB | 16 GB |
Q5_K_M | 5.4 GB | 6.7 GB | 8.1 GB | 8 GB | 16 GB |
Q6_K | 6.3 GB | 7.5 GB | 8.9 GB | 12 GB | 16 GB |
Q8_0 | 8.1 GB | 9.4 GB | 10.8 GB | 12 GB | 16 GB |
F16 | 15.2 GB | 16.5 GB | 17.9 GB | 24 GB | 24 GB* |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.