Reflection's first open-weight model, Beam, has 501 billion parameters, with weights due this month
Reflection says Beam matches GLM 5.2 on reasoning with a third to a quarter of the compute, but DeepSeek V4.1 Flash beats it on Reflection's own coding table.
On October 5, 2026, Reflection announced Beam, its first open-weight model, and said it'll release the weights later this month under an Apache 2.0 license. Beam is a Mixture-of-Experts model with 501 billion parameters, 23 billion of them active at a time, built for coding and agentic work.
You can't download it yet. Beam is still in final red-teaming and evaluations, and an early version is going to a select group of users through a waitlist.
- Maker
- Reflection
- Size
- 501 billion total parameters, 23 billion active
- Context length
- 1M tokens
- Pretraining
- 23.8 trillion tokens
- Weights
- Apache 2.0 license, later this month
- Access now
- waitlist
How it compares
Reflection pitches Beam on efficiency more than raw scores. It says Beam is competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, and it names Kimi K3 as still ahead on raw capability.
On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute.
Reflection's own table is more mixed than that sentence, though. On Terminal Bench v2.1, Beam scores 80.1, against 81.0 for GLM 5.2 and 90.6 for DeepSeek V4.1 Flash. On DeepSWE v1.1 it's 44.4, level with GLM 5.2's 44.0 but far behind DeepSeek V4.1 Flash at 74.2. It's ahead of Nemotron 3 Ultra on most rows where both have a score (SWEBench Verified, 80.9 to 70.7).
The compute comparison is Reflection's own estimate. It counts active parameters times generated tokens and leaves out prompt prefill, so I'd wait for independent runs once the weights are out.
The training run
Most of the post is about reinforcement learning. Reflection ran 10.5K NVIDIA GB300 GPUs for four weeks and generated more than 100 million rollouts (attempts at a task, each one graded), drawing on nearly one million training environments.
We believe this is one of the largest scale RL runs conducted by any open lab to date.
It says capability kept rising as RL compute went up, with no sign of a plateau. Beam's browsing also improved during a phase of training that had no browsing tasks in it. Given web access, the model learned to search for and query other large language models, and to use OCR APIs to read documents.
Beam is text-only. Reflection's demos have it building a live NYC subway map from public MTA data and writing a fine-tuning notebook for the smallest Gemma-4 model.
Safety
Reflection trained a separate safety and alignment model from the same checkpoint and merged the two through distillation. It says it'll publish its safety evaluation results in the technical report, and open-source the safety evaluations it built internally.
The weights, technical report and model card are due later this month.
More on Open weights
- Alibaba says Qwen ran a month of self-improvement unattended, and plans a 10-trillion-parameter modelSeptember 23, 2026
- Qwen models explained: Qwen3.8 Max, Flash and 27B, open weights vs API, with release datesSeptember 23, 2026
- Intrinsic open-sourced the robot control stack it uses in real manufacturing deployments, under Apache 2.0September 22, 2026
- ZCode packaged developers' whole .git history for upload, and Z.ai open-sourced the client in responseSeptember 22, 2026