- Can I run it locally?
- Best GPU for local LLMs
Best GPU for Local LLMs: What Each VRAM Tier Runs
Memory decides which models a machine can run. The weights and the context cache have to fit in memory the GPU can reach, or the runtime pushes layers out to the CPU and every token waits for them. So this page starts from the hardware: each card, Mac and mini PC on sale now, and which of the open models we track fit it, from the same file sizes and arithmetic as our VRAM calculator.
How much VRAM do I need?
Qwen3.8 27B makes a fair yardstick, with 6.3 million downloads of Unsloth's files in 30 days. Its 4-bit file is 16.5 GB, and with a 32K-token context and llama.cpp's buffers it needs 19.8 GB. Take the context to 64K and it's 21.9 GB. The cache (the keys and values a model keeps for every token it has read) is the part that varies, and on an older design like Llama 3.3 70B the full 128K costs 42.9 GB on its own.
The table counts memory the way each machine hands it to the GPU. A graphics card's VRAM is all usable. A Mac gives its GPU two thirds of unified memory up to 32 GB and three quarters above that, by default. AMD lets the Ryzen AI Max+ 395 give its graphics up to 96 GB of 128.
| Machine | Fast memory | Models that fit | Biggest that fits |
|---|---|---|---|
| 16 GB card | 17.2 GB | 28 of 66 | DeepSeek R1 Distill Qwen 14B Q4_K_M |
| 24 GB card | 25.8 GB | 45 of 66 | Qwen AgentWorld 35B A3B UD-Q4_K_M |
| 32 GB card | 34.4 GB | 47 of 66 | DeepSeek R1 Distill Qwen 32B Q4_K_M |
| 48 GB card | 51.5 GB | 47 of 66 | DeepSeek R1 Distill Qwen 32B Q4_K_M |
| 96 GB card | 103 GB | 52 of 66 | Llama 4 Scout 17B 16E Q4_K_M |
| 64 GB Mac | 51.5 GB | 47 of 66 | DeepSeek R1 Distill Qwen 32B Q4_K_M |
| 96 GB Mac | 77.3 GB | 52 of 66 | Llama 4 Scout 17B 16E Q4_K_M |
| 128 GB Mac | 103 GB | 52 of 66 | Llama 4 Scout 17B 16E Q4_K_M |
| 256 GB Mac | 206 GB | 57 of 66 | GLM-5.3 Flash UD-Q4_K_XL |
| 512 GB Mac | 412 GB | 58 of 66 | DeepSeek R1 Q4_K_M |
| Ryzen AI Max+ 395 mini PC | 103 GB | 52 of 66 | Llama 4 Scout 17B 16E Q4_K_M |
| DGX Spark | 133 GB | 54 of 66 | Step 3.7 Flash UD-Q4_K_M |
Each model at its common 4-bit file with a 32K context, counted against the memory the GPU can use: all of a card's VRAM, the default GPU share of a Mac, the graphics allowance AMD sets for Ryzen AI Max.
16, 24 and 32 GB graphics cards: RTX 5060 Ti to RTX 5090
The 16 GB GeForce cards hold exactly the same models, since capacity is the whole test for fitting. What changes as you go up the range is bandwidth, how fast the card reads its own memory: 448 GB/s on the RTX 5060 Ti, 896 on the 5070 Ti and 960 on the 5080, per NVIDIA's comparison table. Each token a model writes reads the weights it uses, so speed follows bandwidth. llama.cpp's Apple benchmark thread shows how closely: a 7B model at 4-bit writes about 14 tokens a second on an M1 with 68 GB/s, and about 94 on an M2 Ultra with 800.
There's no 24 GB GeForce card in the current range. NVIDIA goes from 16 GB on the RTX 5080 to 32 GB on the RTX 5090, so the 24 GB row is AMD's Radeon RX 7900 XTX, a 2022 card, and Intel's Arc Pro B60. llama.cpp runs on both (HIP for AMD, SYCL for Intel, Vulkan for either, from its list of backends). The B60 reads memory at 456 GB/s, under half the 7900 XTX's 960.
16 GB graphics card
28 of the 66 models we track fit a 16 GB card whole at their common 4-bit file with a 32K context. The most downloaded of them is Gemma 4 12B, and the biggest is DeepSeek R1 Distill Qwen 14B at Q4_K_M (16.2 GB).
- DeepSeek R1 Distill Qwen 14B Q4_K_M, 16.2 GB
- gpt-oss-20b Q4_K_M, 13.4 GB
- Gemma 3 12B Q4_K_M, 11.0 GB
- DeepSeek R1 0528 Qwen3 8B Q4_K_M, 10.7 GB
- DeepSeek R1 Distill Llama 8B Q4_K_M, 10.0 GB
- Gemma 4 12B Q4_K_M, 8.9 GB
- GeForce RTX 5060 Ti 16 GB 16 GB GDDR7, 448 GB/s (maker specs)
- GeForce RTX 5070 Ti 16 GB GDDR7, 896 GB/s (maker specs)
- GeForce RTX 5080 16 GB GDDR7, 960 GB/s (maker specs)
24 GB graphics card
Stepping up from a 16 GB card adds 17 models at 4-bit, 45 of 66 in all. The biggest is now Qwen AgentWorld 35B A3B at UD-Q4_K_M (23.9 GB).
- Qwen AgentWorld 35B A3B UD-Q4_K_M, 23.9 GB
- Qwen3.6 35B A3B UD-Q4_K_M, 23.9 GB
- Ornith 1.0 35B UD-Q4_K_M, 23.9 GB
- Qwen3.5 35B A3B Q4_K_M, 23.8 GB
- Qwen3 Coder 30B A3B Q4_K_M, 22.6 GB
- Qwen3 30B A3B 2507 Q4_K_M, 22.6 GB
- Radeon RX 7900 XTX 24 GB GDDR6, 960 GB/s (maker specs)
- Intel Arc Pro B60 24 GB GDDR6, 456 GB/s (maker specs)
32 GB graphics card
A 32 GB card fits 47 of the 66 at 4-bit with 32K of context, 2 more than a 24 GB card, and the biggest of them is DeepSeek R1 Distill Qwen 32B at Q4_K_M (29.3 GB).
- DeepSeek R1 Distill Qwen 32B Q4_K_M, 29.3 GB
- Nemotron 3.5 Lightning 30B A3B UD-Q4_K_M, 26.3 GB
- GeForce RTX 5090 32 GB GDDR7, 1,792 GB/s (maker specs)
- Radeon AI PRO R9700 32 GB GDDR6, 640 GB/s (maker specs)
48 and 96 GB workstation cards: RTX PRO 5000 and RTX PRO 6000
NVIDIA's Blackwell workstation cards are where one card gets past 32 GB. The RTX PRO 5000 has 48 GB of GDDR7 with error correction, and the RTX PRO 6000 has 96 GB at 1,792 GB/s, the RTX 5090's bandwidth with three times the memory (from NVIDIA's spec pages, linked below).
48 GB workstation card
47 of the 66 models we track fit a 48 GB card whole at their common 4-bit file with a 32K context. The most downloaded of them is Qwen3 Coder 30B A3B, and the biggest is DeepSeek R1 Distill Qwen 32B at Q4_K_M (29.3 GB).
- DeepSeek R1 Distill Qwen 32B Q4_K_M, 29.3 GB
- Nemotron 3.5 Lightning 30B A3B UD-Q4_K_M, 26.3 GB
- Qwen AgentWorld 35B A3B UD-Q4_K_M, 23.9 GB
- Qwen3.6 35B A3B UD-Q4_K_M, 23.9 GB
- Ornith 1.0 35B UD-Q4_K_M, 23.9 GB
- Qwen3.5 35B A3B Q4_K_M, 23.8 GB
- RTX PRO 5000 Blackwell 48 GB GDDR7 ECC, 1,344 GB/s (maker specs)
96 GB workstation card
A 96 GB card fits 52 of the 66 at 4-bit with 32K of context, 5 more than a 48 GB card, and the biggest of them is Llama 4 Scout 17B 16E at Q4_K_M (69.2 GB).
- Llama 4 Scout 17B 16E Q4_K_M, 69.2 GB
- gpt-oss-120b Q4_K_M, 64.9 GB
- Llama 3.3 70B Q4_K_M, 54.0 GB
- DeepSeek R1 Distill Llama 70B Q4_K_M, 54.0 GB
- Kimi Dev 72B Q4_0, 52.9 GB
- RTX PRO 6000 Blackwell Workstation Edition 96 GB GDDR7 ECC, 1,792 GB/s (maker specs)
Mac mini and Mac Studio for local LLMs
A Mac's memory is shared between its CPU and GPU, which is how a small desktop holds models no GeForce card can. The rows count the GPU's default share, 48 GB of a 64 GB Mac (per llama.cpp discussion 2182), and sudo sysctl iogpu.wired_limit_mb raises it if you need more.
The Mac mini stops at 64 GB, with the M5 Pro. Past that it's the Mac Studio, which Apple sells with the M5 Max up to 128 GB and the M5 Ultra up to 512 GB, all of it built to order at Apple above 96 GB (Apple's specs). The trade is speed: the 40-core M5 Max reads memory at 614 GB/s and the M5 Ultra at 1.2 TB/s, against 1,792 GB/s on an RTX 5090.
Mac with 64 GB of unified memory
47 of the 66 models we track fit a 64 GB Mac whole at their common 4-bit file with a 32K context. The most downloaded of them is Qwen3 Coder 30B A3B, and the biggest is DeepSeek R1 Distill Qwen 32B at Q4_K_M (29.3 GB).
- DeepSeek R1 Distill Qwen 32B Q4_K_M, 29.3 GB
- Nemotron 3.5 Lightning 30B A3B UD-Q4_K_M, 26.3 GB
- Qwen AgentWorld 35B A3B UD-Q4_K_M, 23.9 GB
- Qwen3.6 35B A3B UD-Q4_K_M, 23.9 GB
- Ornith 1.0 35B UD-Q4_K_M, 23.9 GB
- Qwen3.5 35B A3B Q4_K_M, 23.8 GB
- Mac mini (M5 Pro, 64 GB) 64 GB unified memory (maker specs), $3,199 at Apple (October 1, 2026)
- Mac Studio (M5 Max, 64 GB) 64 GB unified memory, 614 GB/s (maker specs), $3,799 at Apple (October 1, 2026)
Mac with 96 GB of unified memory
Stepping up from a 64 GB Mac adds 5 models at 4-bit, 52 of 66 in all. The biggest is now Llama 4 Scout 17B 16E at Q4_K_M (69.2 GB).
- Llama 4 Scout 17B 16E Q4_K_M, 69.2 GB
- gpt-oss-120b Q4_K_M, 64.9 GB
- Llama 3.3 70B Q4_K_M, 54.0 GB
- DeepSeek R1 Distill Llama 70B Q4_K_M, 54.0 GB
- Kimi Dev 72B Q4_0, 52.9 GB
- Mac Studio (M5 Ultra, 96 GB) 96 GB unified memory, 1,200 GB/s (maker specs), $5,499 at Apple (October 1, 2026)
Mac with 128 GB of unified memory
At 4-bit with 32K of context a 128 GB Mac fits the same 52 models as a 96 GB Mac. The extra memory buys a longer context or a bigger file of the same model.
- Mac Studio (M5 Max, 128 GB) 128 GB unified memory, 614 GB/s (maker specs)
- MacBook Pro 14 (M5 Max, 128 GB) 128 GB unified memory, 614 GB/s (maker specs)
Mac Studio with 256 GB
Stepping up from a 128 GB Mac adds 5 models at 4-bit, 57 of 66 in all. The biggest is now GLM-5.3 Flash at UD-Q4_K_XL (201 GB).
- GLM-5.3 Flash UD-Q4_K_XL, 201 GB
- Inkling Small UD-Q4_K_M, 165 GB
- DeepSeek V4 Flash 0731 UD-Q4_K_XL, 156 GB
- Step 3.7 Flash UD-Q4_K_M, 125 GB
- Qwen3.8 Flash Next UD-Q4_K_XL, 113 GB
- Mac Studio (M5 Ultra, 256 GB) 256 GB unified memory, 1,200 GB/s (maker specs)
Mac Studio with 512 GB
A 512 GB Mac fits 58 of the 66 at 4-bit with 32K of context, 1 more than a 256 GB Mac, and the biggest of them is DeepSeek R1 at Q4_K_M (407 GB).
- DeepSeek R1 Q4_K_M, 407 GB
- Mac Studio (M5 Ultra, 512 GB) 512 GB unified memory, 1,200 GB/s (maker specs)
Best mini PC for local LLMs: Ryzen AI Max+ 395 and DGX Spark
AMD's Ryzen AI Max+ 395 carries 128 GB of LPDDR5X on a 256-bit bus at 256 GB/s, and lets the graphics take up to 96 GB of it (AMD's figures), which is the number we count. GMKtec, Beelink and HP all build small desktops around it with the full 128 GB soldered on.
NVIDIA's DGX Spark is the same idea on its GB10 chip: 128 GB of unified memory at 273 GB/s, and ASUS sells a version of it called the Ascent GX10. Both kinds of box hold models a 32 GB card can't, at about a seventh of an RTX 5090's bandwidth, so expect big models to load fine and then write slowly.
Ryzen AI Max+ 395 mini PC with 128 GB
52 of the 66 models we track fit a Ryzen AI Max+ 395 mini PC whole at their common 4-bit file with a 32K context. The most downloaded of them is Qwen3 Coder 30B A3B, and the biggest is Llama 4 Scout 17B 16E at Q4_K_M (69.2 GB).
- Llama 4 Scout 17B 16E Q4_K_M, 69.2 GB
- gpt-oss-120b Q4_K_M, 64.9 GB
- Llama 3.3 70B Q4_K_M, 54.0 GB
- DeepSeek R1 Distill Llama 70B Q4_K_M, 54.0 GB
- Kimi Dev 72B Q4_0, 52.9 GB
- DeepSeek R1 Distill Qwen 32B Q4_K_M, 29.3 GB
- GMKtec EVO-X2 128 GB LPDDR5X unified memory, 256 GB/s (maker specs)
- Beelink GTR9 Pro 128 GB LPDDR5X unified memory, 256 GB/s (maker specs)
- HP Z2 Mini G1a 128 GB LPDDR5X ECC unified memory, 256 GB/s (maker specs)
DGX Spark or another GB10 box, 128 GB
54 of the 66 models we track fit a DGX Spark whole at their common 4-bit file with a 32K context. The most downloaded of them is Qwen3 Coder 30B A3B, and the biggest is Step 3.7 Flash at UD-Q4_K_M (125 GB).
- Step 3.7 Flash UD-Q4_K_M, 125 GB
- Qwen3.8 Flash Next UD-Q4_K_XL, 113 GB
- Llama 4 Scout 17B 16E Q4_K_M, 69.2 GB
- gpt-oss-120b Q4_K_M, 64.9 GB
- Llama 3.3 70B Q4_K_M, 54.0 GB
- DeepSeek R1 Distill Llama 70B Q4_K_M, 54.0 GB
- NVIDIA DGX Spark 128 GB LPDDR5X unified memory, 273 GB/s (maker specs)
- ASUS Ascent GX10 128 GB LPDDR5X unified memory, 273 GB/s (maker specs)
Ollama hardware requirements
Ollama runs the same kind of GGUF file, so the arithmetic holds, with two defaults worth knowing. A plain tag like gemma3:12b is the 4-bit Q4_K_M build: on Ollama's tags page both are 8.1 GB. And the context is set from your VRAM. Ollama's context length page gives 4K tokens under 24 GiB, 32K from 24 to 48 GiB and 256K at 48 GiB and up. Its FAQ still says 4,096 tokens for everyone, so the two pages disagree.
A model too big for the card still loads: Ollama splits the layers between the GPU and system memory, and ollama ps shows the split (48%/52% CPU/GPU, say). The cache is f16 unless you set OLLAMA_KV_CACHE_TYPE; q8_0 there takes about half the memory. The table reruns our numbers at Ollama's default context for each card.
Under those defaults a 48 GB card fits 45 of them and a 32 GB card fits 47. The bigger card gets the longer default context, 256K against 32K, and the cache eats the difference until you set OLLAMA_CONTEXT_LENGTH lower.
| Card | Ollama default context | Fit at 4-bit | Biggest that fits |
|---|---|---|---|
| 16 GB graphics card | 4K | 31 of 66 | Muse Glimmer 30B |
| 24 GB graphics card | 32K | 45 of 66 | Qwen AgentWorld 35B A3B |
| 32 GB graphics card | 32K | 47 of 66 | DeepSeek R1 Distill Qwen 32B |
| 48 GB workstation card | 256K | 45 of 66 | Qwen3 Coder 30B A3B |
| 96 GB workstation card | 256K | 52 of 66 | Llama 3.3 70B |
How the numbers are worked out
Total memory is three parts added together. Weights are the exact size of the GGUF file on Hugging Face, all parts summed when a quant is split. The context cache (the KV cache) comes from each model's own config.json:
cache = layers × kv_heads × head_dim × 2 (K and V) × 2 bytes × tokens
Sliding-window layers stop at their window plus llama.cpp's 512-token batch. Linear-attention and Mamba layers keep a fixed state instead, which we add once. Latent-attention models (DeepSeek, Kimi, GLM-5) cache kv_lora_rank + rope dimensions per layer, the way llama.cpp does. Overhead is our allowance: 0.5 GB for the GPU context plus the output buffer, 512 × vocabulary × 4 bytes.
A "24 GB" card holds 24 GiB, which is 25.8 GB in the decimal units file sizes use, and we count it that way. Macs get the share Metal allows by default, two thirds of unified memory up to 32 GB and three quarters above (llama.cpp discussion 2182); CPU-only machines and the DGX Spark keep 4 GB back for the operating system. All of it assumes llama.cpp's defaults: an f16 cache and one conversation at a time. Quantizing the cache to q8_0 halves it.
Files are Unsloth's GGUF uploads, the most downloaded quantizations on Hugging Face, and the list refreshes daily from Hugging Face's API. A model joins when its repo passes 20,000 downloads in 30 days. Last checked September 30, 2026.