Kimi VRAM requirements: can I run Kimi K3 locally?
Kimi K3 has 2.78 trillion parameters, and even its smallest file is UD-Q1_0 (466 GB). Here's the one desktop that holds it, and what the rest of the Kimi family needs.
Check your machine
Won't fit 1,511 GB
Weights, cache and overhead add up to more than the card and the RAM together.
- Weights 1,509 GB
- Context cache 1.4 GB
- Overhead 0.8 GB
- Fast memory
- With system RAM
What we found
Kimi K3 won't run on a Mac mini or a DGX Spark. Unsloth's smallest file, UD-Q1_0 (466 GB), is more than three times the Spark's 128 GB. The one single machine it fits is a 512 GB Mac Studio, running that smallest file with the GPU memory limit raised.
The cache is the one part of Kimi K3 that's modest. Only 24 of its 93 layers keep one, and they store a compressed latent, so a full million-token context adds 29.5 GB. Next to the weights that barely registers.
Kimi K2.7 Code and K2.6 share one size, 1.03 trillion parameters. Their smallest files fit a 512 GB Mac without touching the GPU limit.
Kimi Dev 72B is the odd one out: a 72B dense model built on Qwen2.5, so it behaves like a 70B Llama for memory, 52.9 GB with a 32K context.
Kimi K3
2.78 trillion parameters · mixture of experts, 16 of 896 active · up to 1M context · 436 thousand downloads in 30 days
Kimi K3 needs about 1,511 GB at UD-Q4_K_XL with a 32K-token context. The file itself is 1,509 GB, and the cache for 32K tokens adds 1.4 GB. That's past a 512 GB Mac. The smallest file, UD-Q1_0 at 466 GB, fits a 512 GB Mac, with the GPU memory limit raised. It routes each token through 16 of 896 experts, so it runs faster than its size suggests.
Only 24 of its 93 layers keep a growing cache (the rest carry a fixed-size state), so the full 1M context costs 29.5 GB of cache.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q1_0 | 466 GB | 468 GB | 469 GB | none | 512 GB* |
UD-Q2_K_XL | 861 GB | 863 GB | 863 GB | none | none |
UD-Q4_K_XL | 1,509 GB | 1,510 GB | 1,511 GB | none | none |
UD-Q8_K_XL | 1,561 GB | 1,563 GB | 1,563 GB | none | none |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Kimi K2.7 Code
1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 256K context · 356 thousand downloads in 30 days
At UD-Q4_K_XL, the Kimi K2.7 Code file is 584 GB, and with 32K tokens of context and runtime overhead it comes to about 587 GB. That's past a 512 GB Mac. The smallest file, UD-IQ1_M at 304 GB, fits a 512 GB Mac. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.
It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 18.4 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_M | 304 GB | 305 GB | 307 GB | none | 512 GB |
UD-Q2_K_XL | 339 GB | 341 GB | 343 GB | none | 512 GB |
UD-Q3_K_XL | 464 GB | 465 GB | 467 GB | none | 512 GB* |
UD-Q4_K_XL | 584 GB | 585 GB | 587 GB | none | none |
UD-Q8_K_XL | 595 GB | 596 GB | 598 GB | none | none |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Kimi Dev 72B
72.7 billion parameters · up to 128K context · 23 thousand downloads in 30 days
Kimi Dev 72B needs about 52.9 GB at Q4_0 with a 32K-token context. The file itself is 41.4 GB, and the cache for 32K tokens adds 10.7 GB. That fits a RTX PRO 6000 (96 GB).
Every token of context adds 320 KB of cache, so the full 128K would put 42.9 GB on top of the weights.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-IQ1_S | 23.0 GB | 26.5 GB | 34.6 GB | 48 GB | 48 GB |
UD-Q2_K_XL | 30.3 GB | 33.8 GB | 41.9 GB | 48 GB | 48 GB* |
UD-Q3_K_XL | 38.3 GB | 41.8 GB | 49.8 GB | 48 GB | 64 GB |
Q4_0 | 41.4 GB | 44.9 GB | 52.9 GB | 96 GB | 64 GB* |
UD-Q5_K_XL | 54.1 GB | 57.6 GB | 65.7 GB | 96 GB | 96 GB |
Q6_K | 64.3 GB | 67.8 GB | 75.9 GB | 96 GB | 96 GB |
Q8_0 | 77.3 GB | 80.8 GB | 88.8 GB | 96 GB | 96 GB* |
BF16 | 145 GB | 149 GB | 157 GB | none | 256 GB |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
Kimi K2.6
1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 256K context · 253 thousand downloads in 30 days
At UD-Q4_K_XL, the Kimi K2.6 file is 584 GB, and with 32K tokens of context and runtime overhead it comes to about 587 GB. That's past a 512 GB Mac. The smallest file, UD-Q2_K_XL at 340 GB, fits a 512 GB Mac. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.
It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 18.4 GB at 256K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-Q2_K_XL | 340 GB | 342 GB | 344 GB | none | 512 GB |
UD-Q4_K_XL | 584 GB | 585 GB | 587 GB | none | none |
UD-Q8_K_XL | 595 GB | 596 GB | 598 GB | none | none |
BF16 | 2,053 GB | 2,055 GB | 2,056 GB | none | none |
Smallest GPU and Mac are for the 32K total.
Sources: GGUF files · config.json
Kimi K2
1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 128K context · 44 thousand downloads in 30 days
Kimi K2 needs about 624 GB at Q4_K_M with a 32K-token context. The file itself is 621 GB, and the cache for 32K tokens adds 2.3 GB. That's past a 512 GB Mac. The smallest file, UD-TQ1_0 at 244 GB, fits a 256 GB Mac, with the GPU memory limit raised. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.
It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 9.2 GB at 128K.
| Quant | File | At 8K | At 32K | Smallest GPU | Smallest Mac |
|---|---|---|---|---|---|
UD-TQ1_0 | 244 GB | 245 GB | 247 GB | none | 256 GB* |
UD-Q2_K_XL | 382 GB | 383 GB | 385 GB | none | 512 GB |
UD-Q3_K_XL | 452 GB | 453 GB | 455 GB | none | 512 GB* |
Q4_K_M | 621 GB | 622 GB | 624 GB | none | none |
Q5_K_M | 728 GB | 730 GB | 731 GB | none | none |
Q6_K | 843 GB | 844 GB | 846 GB | none | none |
Q8_0 | 1,091 GB | 1,092 GB | 1,094 GB | none | none |
BF16 | 2,053 GB | 2,055 GB | 2,056 GB | none | none |
Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.
Sources: GGUF files · config.json
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.