Open-weight AI model news: Qwen, DeepSeek, Mistral and more
Models whose weights anyone can download and run on their own hardware, and the companies that release them, newest first. Loading errors and context limits for those models are here too, since running one yourself is where they tend to show up.
Alibaba says Qwen ran a month of self-improvement unattended, and plans a 10-trillion-parameter model
Eddie Wu's keynote says the Qwen team has made meaningful progress on AI that improves itself, the trend Dario Amodei wants slowed.
Qwen models explained: Qwen3.8 Max, Flash and 27B, open weights vs API, with release dates
Alibaba put its Max-class Qwen out as open weights for the first time, and skipped open weights for Qwen3.7 entirely.
“Error: pull model manifest: file does not exist” in Ollama: the name or tag isn't on the registry
What Ollama means by “Error: pull model manifest: file does not exist”, why it's really a 404 from the model registry, the :cloud suggestion newer versions add, and the /save bug that prints it for a model you already have.
Intrinsic open-sourced the robot control stack it uses in real manufacturing deployments, under Apache 2.0
The GitHub repository carries one line under the license: this is not an officially supported Google product.
ZCode packaged developers' whole .git history for upload, and Z.ai open-sourced the client in response
The rule that allows .git runs ahead of every exclusion in ZCode's filter chain, so neither the 1 MB size cap nor the key filter ever touches what's inside it.
Xiaomi's MiMo-V2.6-Pro tops the open-weights index, but its own table puts Opus 5 ahead
Xiaomi trained it with one reinforcement learning run across coding, agents, vision and cybersecurity, and put the weights on Hugging Face under MIT.
Qwen open-sourced Qwen-Image-2.1's weights, but the license is research-only where Qwen-Image was Apache 2.0
The 7B model makes transparent images straight from a prompt, and it takes up to 10 reference pictures for a single edit.
llama.cpp unknown model architecture: 'qwen4exp': what it means and how to fix it
What llama.cpp means by unknown model architecture: 'qwen4exp', which pull request added Qwen3.8-Flash-Next support and when, the Ollama release that carries it, and the load failures that look similar.
What's the context limit on Qwen3.8-27B, and how do you push it to 1M?
The million-token number needs a config change, and Qwen says leaving that change on hurts short prompts.
llama.cpp hangs loading Qwen3.8-Flash-Next on ROCm: what the last log line means
Why llama-server stops partway through loading Qwen3.8-Flash-Next on llama.cpp's HIP backend, why the hang isn't an out-of-memory failure, what sets it off, and what loads the same file instead.
Firefox's AI assistant now runs on Mistral, and Mozilla left the model picker in place.
Two companies that call themselves open source advocates are selling a browser deal as a brake on closed-model AI.
What does the DeepSeek V4 Pro API cost, and what are its rate limits?
Peak covers seven hours a day, Monday to Friday, so most of the week bills at half the peak rate.
How many DeepSeek V4 Flash tokens do you get on Ollama Cloud Pro?
Ollama doesn't publish a token allowance for the plan, so the answer is $60 divided by a per-million rate.
DeepSeek's V4.1 Flash edged past GPT-6 Astra on one independent test
It's 13 points behind Astra on the overall index, and DeepSeek has already backed off retiring V4 Pro in its favor.




