ThinkFacility
  1. Can I run it locally?
  2. Kimi

Kimi VRAM requirements: can I run Kimi K3 locally?

Kimi K3 has 2.78 trillion parameters, and even its smallest file is UD-Q1_0 (466 GB). Here's the one desktop that holds it, and what the rest of the Kimi family needs.

Check your machine

Won't fit 1,511 GB

Weights, cache and overhead add up to more than the card and the RAM together.

  • Weights 1,509 GB
  • Context cache 1.4 GB
  • Overhead 0.8 GB
  • Fast memory
  • With system RAM

What we found

Kimi K3 won't run on a Mac mini or a DGX Spark. Unsloth's smallest file, UD-Q1_0 (466 GB), is more than three times the Spark's 128 GB. The one single machine it fits is a 512 GB Mac Studio, running that smallest file with the GPU memory limit raised.

The cache is the one part of Kimi K3 that's modest. Only 24 of its 93 layers keep one, and they store a compressed latent, so a full million-token context adds 29.5 GB. Next to the weights that barely registers.

Kimi K2.7 Code and K2.6 share one size, 1.03 trillion parameters. Their smallest files fit a 512 GB Mac without touching the GPU limit.

Kimi Dev 72B is the odd one out: a 72B dense model built on Qwen2.5, so it behaves like a 70B Llama for memory, 52.9 GB with a 32K context.

Kimi K3

2.78 trillion parameters · mixture of experts, 16 of 896 active · up to 1M context · 436 thousand downloads in 30 days

Kimi K3 needs about 1,511 GB at UD-Q4_K_XL with a 32K-token context. The file itself is 1,509 GB, and the cache for 32K tokens adds 1.4 GB. That's past a 512 GB Mac. The smallest file, UD-Q1_0 at 466 GB, fits a 512 GB Mac, with the GPU memory limit raised. It routes each token through 16 of 896 experts, so it runs faster than its size suggests.

Only 24 of its 93 layers keep a growing cache (the rest carry a fixed-size state), so the full 1M context costs 29.5 GB of cache.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q1_0466 GB468 GB469 GBnone512 GB*
UD-Q2_K_XL861 GB863 GB863 GBnonenone
UD-Q4_K_XL1,509 GB1,510 GB1,511 GBnonenone
UD-Q8_K_XL1,561 GB1,563 GB1,563 GBnonenone

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Kimi K2.7 Code

1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 256K context · 356 thousand downloads in 30 days

At UD-Q4_K_XL, the Kimi K2.7 Code file is 584 GB, and with 32K tokens of context and runtime overhead it comes to about 587 GB. That's past a 512 GB Mac. The smallest file, UD-IQ1_M at 304 GB, fits a 512 GB Mac. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.

It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 18.4 GB at 256K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_M304 GB305 GB307 GBnone512 GB
UD-Q2_K_XL339 GB341 GB343 GBnone512 GB
UD-Q3_K_XL464 GB465 GB467 GBnone512 GB*
UD-Q4_K_XL584 GB585 GB587 GBnonenone
UD-Q8_K_XL595 GB596 GB598 GBnonenone

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Kimi Dev 72B

72.7 billion parameters · up to 128K context · 23 thousand downloads in 30 days

Kimi Dev 72B needs about 52.9 GB at Q4_0 with a 32K-token context. The file itself is 41.4 GB, and the cache for 32K tokens adds 10.7 GB. That fits a RTX PRO 6000 (96 GB).

Every token of context adds 320 KB of cache, so the full 128K would put 42.9 GB on top of the weights.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-IQ1_S23.0 GB26.5 GB34.6 GB48 GB48 GB
UD-Q2_K_XL30.3 GB33.8 GB41.9 GB48 GB48 GB*
UD-Q3_K_XL38.3 GB41.8 GB49.8 GB48 GB64 GB
Q4_041.4 GB44.9 GB52.9 GB96 GB64 GB*
UD-Q5_K_XL54.1 GB57.6 GB65.7 GB96 GB96 GB
Q6_K64.3 GB67.8 GB75.9 GB96 GB96 GB
Q8_077.3 GB80.8 GB88.8 GB96 GB96 GB*
BF16145 GB149 GB157 GBnone256 GB

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

Kimi K2.6

1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 256K context · 253 thousand downloads in 30 days

At UD-Q4_K_XL, the Kimi K2.6 file is 584 GB, and with 32K tokens of context and runtime overhead it comes to about 587 GB. That's past a 512 GB Mac. The smallest file, UD-Q2_K_XL at 340 GB, fits a 512 GB Mac. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.

It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 18.4 GB at 256K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-Q2_K_XL340 GB342 GB344 GBnone512 GB
UD-Q4_K_XL584 GB585 GB587 GBnonenone
UD-Q8_K_XL595 GB596 GB598 GBnonenone
BF162,053 GB2,055 GB2,056 GBnonenone

Smallest GPU and Mac are for the 32K total.

Sources: GGUF files · config.json

Kimi K2

1.03 trillion parameters · mixture of experts, 8 of 384 active · up to 128K context · 44 thousand downloads in 30 days

Kimi K2 needs about 624 GB at Q4_K_M with a 32K-token context. The file itself is 621 GB, and the cache for 32K tokens adds 2.3 GB. That's past a 512 GB Mac. The smallest file, UD-TQ1_0 at 244 GB, fits a 256 GB Mac, with the GPU memory limit raised. It routes each token through 8 of 384 experts, so it runs faster than its size suggests.

It caches a compressed latent per token rather than full keys and values, about 69 KB a token, so the cache is small next to the weights: 9.2 GB at 128K.

QuantFileAt 8KAt 32KSmallest GPUSmallest Mac
UD-TQ1_0244 GB245 GB247 GBnone256 GB*
UD-Q2_K_XL382 GB383 GB385 GBnone512 GB
UD-Q3_K_XL452 GB453 GB455 GBnone512 GB*
Q4_K_M621 GB622 GB624 GBnonenone
Q5_K_M728 GB730 GB731 GBnonenone
Q6_K843 GB844 GB846 GBnonenone
Q8_01,091 GB1,092 GB1,094 GBnonenone
BF162,053 GB2,055 GB2,056 GBnonenone

Smallest GPU and Mac are for the 32K total. * with the Mac's GPU memory limit raised.

Sources: GGUF files · config.json

How the numbers are worked out

Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:

cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens

Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.

A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.

Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 25, 2026.