Reflection AI Beam: 501B Open Weight Model, Explained
Reflection AI's Beam packs 501B total parameters, 23B active, trained on 23.8T tokens, and is set for an Apache 2.0 release.

What is Reflection AI’s Beam model?
Beam is a large mixture of experts language model announced by Reflection AI, built with 501 billion total parameters but only 23 billion active per inference step. Reflection says it will release the weights under an Apache 2.0 license later this month, positioning Beam as an open alternative to closed enterprise models from OpenAI and Anthropic, and as a Western competitor to Chinese open weight releases like Kimi K2 and GLM. As of the announcement, the weights were not yet public, so the model’s real-world performance still needs independent verification.
TL;DR
- Beam uses a mixture of experts design, with 501 billion total parameters but only 23 billion active at any given time, which keeps inference costs down while retaining a large overall knowledge base.
- Training ran on 23.8 trillion tokens, filtered down from a much larger raw internet dataset, with Reflection claiming roughly 95% of raw data was discarded and about 1.8 trillion tokens were rescued that standard filters would have missed.
- Pre-training used 6,144 Nvidia GPUs over about four weeks, with a reported 92.3% useful training time and only nine restarts, numbers that suggest a fairly stable, well-engineered run.
- A large reinforcement learning phase followed pre-training, running on 10,500 GPUs for four weeks, generating over 100 million attempts across roughly one million practice environments.
- Beam leads some open model benchmarks but trails on others, outperforming models like Nemotron on tests such as SWE-Bench, while losing to Kimi K2, GLM, and Qwen3 Max on Terminal Bench and other reasoning tasks.
- The weights are not released yet, so every figure so far comes from Reflection AI’s own announcement rather than independent testing.
- Beam arrives alongside a broader industry split, where companies like Cohere push locked-down, enterprise-controlled AI while others bet on open weights as a trust and distribution strategy.
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
How does Beam’s mixture of experts architecture work?
Mixture of experts (MoE) is a model design where not every parameter gets used for every input. Instead of one giant dense network processing every token with its full parameter count, the model is split into many smaller “expert” subnetworks, and a routing mechanism picks a small subset of experts for each token.
For Beam, that means the model has 501 billion parameters in total, but only 23 billion are active during any single inference pass. The practical effect is that Beam behaves computationally more like a 23-billion-parameter model while retaining access to the broader knowledge encoded across all 501 billion parameters. This is the same basic approach used by other large MoE models in the open weight space, including Kimi K2 and GLM, and it’s part of why very large models can still run at manageable cost, whether self-hosted or served through an API.
What does Beam’s training process reveal?
Reflection AI pre-trained Beam on 23.8 trillion tokens using a cluster of 6,144 Nvidia GB300 GPUs, completing the run in under four weeks. The company says it aggressively filtered its raw data, discarding about 95% of what it initially collected. At the same time, it claims standard filtering approaches would have missed around 1.8 trillion tokens of usable data that its own pipeline managed to keep. The implied argument is that data quality, not just data volume, is what separates a competitive model from a merely large one.
On the infrastructure side, Reflection published a pre-training loss curve that it describes as unusually smooth, with only nine restarts across the entire run and 92.3% of total time spent on productive training rather than recovering from failures or instability. Large training runs often show sharp spikes where loss destabilizes, so a clean curve with minimal restarts is being presented as evidence of solid engineering discipline, not just raw compute scale.
How big was the reinforcement learning phase?
After pre-training, Reflection ran a large reinforcement learning (RL) stage, which the company frames as central to Beam’s capabilities. RL works by having the model attempt tasks repeatedly and adjusting based on feedback on which attempts succeeded, similar to a student working through a massive set of practice problems and getting graded on each one.
For Beam, this phase used 10,500 GPUs over four weeks, generating more than 100 million attempts across roughly one million distinct practice environments. Reflection describes this as one of the largest open RL runs by any lab and claims benchmark scores were still improving when the run was stopped, suggesting there may have been more performance left on the table had training continued.
How does Beam compare to Kimi K2, GLM, and Qwen?
Reflection published benchmark comparisons against other open models, including Nemotron, GLM, Qwen3 Max, and Kimi K2. Beam performs well against Nemotron and leads on tests like SWE-Bench (reported at 80.9) and Terminal Bench 2.1 (reported at 80.1).
But against the larger, more established open models, the picture is mixed. On Terminal Bench, GLM scored 81.0 to 81.2, Qwen3 Max scored 86.6, and Kimi K2 scored 88.3, all ahead of Beam’s reported number. Reflection itself acknowledges that Kimi K2 leads on raw capability in several areas, including reasoning and tool use. Reflection’s positioning isn’t that Beam is the single best open model available, it’s that Beam is the most efficient option and the strongest Western-built open weight model on the market, a narrower but still meaningful claim if it holds up.
Is Beam worth paying attention to before the weights are released?
Right now, every number tied to Beam comes from Reflection AI’s own announcement, not from independent benchmarking. That doesn’t mean the figures are wrong, but it does mean they haven’t been stress-tested by outside researchers or developers running the model themselves. The real test will come when the weights are actually published under Apache 2.0 as promised, and the broader community can verify benchmark claims, check inference costs in practice, and see how Beam performs on tasks outside Reflection’s own published test suite.
If the release happens on schedule and the numbers roughly hold, Beam would be a notable entry in the open weight space, particularly as a Western-built alternative to Chinese MoE models like Kimi K2 and GLM. If the release slips or underperforms independent testing, it becomes another example of a lab publishing strong numbers ahead of a model’s actual availability, something that’s become common across the industry as labs compete for attention before weights are downloadable.
Why does Beam’s release matter for the open weight AI market?
Beam’s announcement lands at a moment when major AI labs are under pressure to show their spending translates into revenue. OpenAI and Anthropic dominate enterprise AI while reportedly operating at significant losses, and investors are pushing for returns. That pressure shapes how other companies position themselves. Some, particularly several Chinese labs, now hold back open weight releases for weeks after announcing a model, monetizing early access before opening things up. Others, especially European and non-American companies outside the OpenAI/Anthropic orbit, are pivoting toward enterprise sales built around control, governance, and predictable costs rather than raw model capability.
Reflection AI’s approach runs counter to the enterprise-lockdown trend. By promising an Apache 2.0 release of a 501-billion-parameter model, Reflection is betting that openness itself is a competitive advantage, letting developers and companies run a capable model on their own infrastructure without being tied to a single provider’s API or pricing. Whether that bet pays off depends entirely on whether the weights ship as promised and whether the model holds up once people outside Reflection AI can actually test it.
Frequently Asked Questions
What does “501B total, 23B active” mean for Beam?
It means Beam has 501 billion parameters in total, but its mixture of experts architecture only activates about 23 billion of them for any given input. This keeps compute and memory costs during inference much closer to a 23-billion-parameter model while still drawing on the broader capacity of the full network.
When will Beam’s weights be released?
Reflection AI stated the weights would be released later in the month of the announcement under an Apache 2.0 license. As of the announcement, the weights were not yet public, so this remains a commitment rather than a completed release.
How does Beam compare to Kimi K2?
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
On published benchmarks, Kimi K2 outperforms Beam on tests like Terminal Bench (88.3 versus Beam’s 80.1), and Reflection AI itself acknowledges Kimi K2 leads on raw capability in areas like reasoning and tool use. Beam’s pitch is efficiency and being a leading Western-built open weight option, not outright superiority over every competing open model.
What license will Beam use?
Reflection AI says Beam will be released under an Apache 2.0 license, a permissive open source license that allows commercial use, modification, and redistribution with minimal restrictions.
Why does data filtering matter so much in Beam’s training?
Reflection AI reports discarding about 95% of raw internet data during cleaning while specifically preserving around 1.8 trillion tokens that standard filtering methods would typically miss. The implicit claim is that careful data curation, rather than sheer token volume, is a major driver of model quality.





