Error: 500 Internal Server Error: unable to load model:
The message
Error: 500 Internal Server Error: unable to load model:What it means
Ollama handed the model file to its bundled copy of llama.cpp and llama.cpp gave nothing back, so Ollama reported the blob's path and stopped. It almost always means that copy of llama.cpp doesn't know the model's architecture yet.
What to do
Update Ollama: the line comes from 0.24.0 and earlier, and split Hugging Face models were being fixed in the 0.30.0 release candidates. If it's a community upload, try the library version of the same model.
The pull finishes, the digest checks out, and then ollama run fails straight away with a path into your models folder:
Error: 500 Internal Server Error: unable to load model: /Users/paul/.ollama/models/blobs/sha256-2862bb0d66a0fe2799de099ee5fdb242449e16f1eaee1091677112d7ca2859c8
That one is from #15515, filed in April 2026. On Windows the path starts with a drive letter, and through the OpenAI-compatible endpoint the same text comes back inside {"error":{"message":"unable to load model: ...","type":"api_error"}}. As of October 1, 2026 there were 184 issues in the ollama/ollama repo containing the phrase, and in the recent ones it's mostly qwen3.5, gemma4 and other newly released models.
What the message comes from
It's one line in llama/llama.go, the Go wrapper around the copy of llama.cpp that Ollama ships. When Ollama can't run a model on its own engine, it falls back to llama.cpp (Ollama's log calls it "compatibility mode"). If llama.cpp's loader returns nothing for the file, the wrapper reports unable to load model: plus the blob path. llama.cpp's own reason goes to the server log, so the 500 you see carries no detail. One reporter in #14575 found their server.log empty as well.
We checked which releases contain the line. It's there in every tag we looked at up to 0.24.0 (May 14, 2026) and gone from 0.30.0 onward, whose release notes promise improved compatibility using llama.cpp. The latest release, 0.35.0, doesn't have the string at all. If you're seeing it, you're on 0.24 or older.
Why the model wouldn't load
Collaborator rick-github explained the qwen3.5 case in #14575 (March 3, 2026). Models from Hugging Face come split, with the text weights and the vision part in separate GGUF files, and split models have to run on the llama.cpp engine. "If the llama.cpp engine doesn't support the architecture, the model won't run." At that point Ollama's bundled llama.cpp didn't know qwen35, qwen35moe or gemma4.
Other threads land in the same place. Whisper uploads fail because, in his words, "Ollama does not currently support whisper." In #16164 a gemma-4 variant pulled from a user's page on ollama.com failed the same way, and his answer was that those are user-uploaded models. The main library is the part the team tries to keep working.
What fixed it
Updating. rick-github wrote on May 15, 2026 that split qwen3.5 support was coming in pull request #16031, then in testing as a 0.30.0 release candidate, and showed a Hugging Face Qwen3.5 GGUF and a gemma-4 GGUF both running on it. Run ollama -v first, since a lot of reports came from people several versions behind.
There's a catch for hf.co/... names. Ollama member dhiltgen wrote in June 2026 that split models "are supported, but only via the create flow", because pulling straight from the hf registry can't glue the GGUFs into one. The workaround he posted in #5245 is to download the files from Hugging Face yourself and build the model with ollama create.
Stuck on an old version? Before 0.30 the suggestion was to drop the vision file. rick-github posted the steps in #14503: keep only the text blob's FROM line from ollama show --modelfile, add the template from the library version of the same model (Hugging Face uploads usually lack a correct one, he noted), then ollama create a new name. You lose image input.
If the architecture is something Ollama doesn't run at all, such as whisper, no update will change it. Pick a different model.
Other lines the same feature prints
Match yours against these if the one at the top of the page is not quite it. They come from the same code and mean related things.
500 Internal Server Error: unable to load modelunable to load model: /Users/paul/.ollama/models/blobs/sha256-2862bb0d66a0fe2799de099ee5fdb242449e16f1eaee1091677112d7ca2859c8{"error":{"message":"unable to load model: /Users/paul/.ollama/models/blobs/sha256-2862bb0d66a0fe2799de099ee5fdb242449e16f1eaee1091677112d7ca2859c8","type":"api_error","param":null,"code":null}}