On October 5, 2026, Reflection AI announced Beam — its first open-weight frontier model. The company is positioning the model as the West's answer to China's open models: DeepSeek, Qwen and Z.ai. Beam is a text-only mixture-of-experts (MoE) architecture: 501 billion parameters in total, 23 billion active; pre-trained on 23.8 trillion tokens with a 1-million-token context window. The model was trained with high-compute reinforcement learning for coding, reasoning and agentic tasks, and will be distributed through hyperscalers and neoclouds. The goal of the model is to deliver more intelligence per token, making enterprise coding and agentic workloads cheaper.
Model Overview
Beam is a fully text-based model: it does not process images or other modalities directly, but can work with them when they are described in text. The architecture is built on a roaming MoE design: interconnected local and global attention layers, fine-grained routed experts and load-balancing mechanisms that ensure even load distribution. Pre-training used 23.8 trillion diverse, curated tokens — from web sources and licensed private data sets. According to the company, the final Beam Base model matched or exceeded the results of existing open base models of comparable size.
High-Compute Reinforcement Learning
The main differentiator of Beam is large-scale reinforcement learning aimed at developing reasoning and agentic capabilities. According to the Reflection blog, the RL process generated more than 100 million rollouts over four weeks on 10,500 NVIDIA GB300 GPUs, with roughly 1.3 billion sandboxes used for evaluation. The company calls this scale "one of the largest RL runs among open labs." During training, a pool of nearly one million high-quality coding, agentic and STEM environments was built; task quality was tested at every stage.
The "reasoning effort" parameter controls the model's depth of thought: a low setting produces a short answer, while a high setting allows longer reasoning on complex tasks. The company says that at the start of RL the model learned to find more efficient solutions with shorter reasoning, and later, as stronger agentic capabilities developed, response length grew again — but the additional tokens delivered new efficiency gains.
Benchmark Results
The company claims Beam matched Z.ai's GLM-5.2 model on advanced reasoning benchmarks while spending 3–4 times less compute at inference. Published coding benchmark scores: 80.9% on SWE Bench Verified, 80.1% on Terminal Bench 2.1, 77.2% on SWE Bench Pro v2-Hard. Its 44.4% result on DeepSWE v1.1 is nearly identical to GLM-5.2's 44.0%. As the Reflection blog notes, the efficiency gap is even more pronounced against models with over 2 trillion parameters, such as Qwen 3.8-Max.
"Together these efforts delivered competitive open-weight performance alongside inference compute efficiency." — Reflection AI blog
TechCrunch reminds readers that these claims have not yet undergone independent verification. The blog's comparison tables also show models like Kimi K3 and DeepSeek V4.1 Flash leading on some benchmarks — Beam's main advantage is not absolute leadership, but inference efficiency.
Generalizing Capabilities
According to the company, RL training shaped agentic capabilities in the model that extend beyond the tasks it was trained on. For example, despite the absence of web-browsing tasks in the training mix, the model showed consistent gains on browsing metrics. Given web access, Beam independently learned to search for and query other large language models and to use OCR APIs for reading documents. In company demos, the model completed tasks such as building a live New York subway board from public data, creating interactive apps and preparing a fine-tuned notebook for a Text2SQL task on Gemma-4.
Release Plan
The weights, technical report, model card and developer artifacts will be released later this month. At launch, Beam will be distributed through hyperscalers and neoclouds, with integrations with open-source libraries ready at the same time. Beam is currently undergoing final red-team testing and evaluations, and registration for early access is open.
Behind the Company
Reflection AI was founded in 2024 by two former Google DeepMind researchers. The company has raised about $4.7 billion with backing from Nvidia, Sequoia and Lightspeed, reaching a valuation of $25 billion. It has also signed compute agreements worth more than $7 billion with SpaceX and Nebius to secure Nvidia GB300 capacity through 2029.



