The pitch is efficiency, not supremacy

Beam is a text-only mixture-of-experts model: 501 billion total parameters, 23 billion active per token, pretrained on 23.8 trillion tokens, with a 1-million-token context window. For scale, Z.ai's GLM-5.2 carries about 744 billion total parameters with 40 billion active. Reflection says Beam was trained with high-compute reinforcement learning so it reasons in fewer steps, and calls the result a "workhorse model" for enterprises, the public sector, and developers.

The headline number is not a benchmark score. It is a cost ratio. Reflection claims Beam scores comparably to GLM-5.2 on advanced reasoning benchmarks at three to four times less inference compute, measured across coding and agentic tasks. The company frames that as an estimate covering its DeepSWE, Humanity's Last Exam, and Terminal Bench numbers, not a measured dollar cost. That distinction matters, because cost-per-token is the argument the whole company rests on.

Its own chart shows the gap

Reflection published its comparison table, and the honest reading is less flattering than the announcement. On most rows, newer Chinese open models are ahead. Terminal Bench v2.1: Beam 80.1, GLM-5.3 88.2, Kimi K3 88.3. Humanity's Last Exam (no tools): Beam 36.2, Kimi K3 46.9. GPQA Diamond: Beam 90.5, Kimi K3 93.5. The bright spot is SWE-Bench Pro v1, where Beam's 65.5 beats GLM-5.2's 62.1, and Reflection notes Beam outscores Inkling, the open model from Mira Murati's Thinking Machines Lab, on four coding tests where both report scores. The footnote there is that Inkling is multimodal and Beam handles text only.

None of these scores are independently verified. They are company-reported, sourced from Reflection's table with rival numbers pulled from Artificial Analysis and DataCurve. TechCrunch flagged that the claims have not been checked against a public checkpoint, because there is no public checkpoint to check.

"Open-weight" is a schedule, not a download

Beam was announced Monday, but only a waitlisted group gets an early version for now. The open weights, a technical report, and a model card are due later this month. Calling it open today is a calendar claim. Two days ago this site wrote that Reflection first had to ship a model. It has now shipped the announcement of one. The model itself is still in the warehouse.

That sequencing is worth noticing because the American-open-model story depends on it. Axios reported over the weekend that a launch was close, and Reflection confirmed the timing. The company was founded in 2024 by Misha Laskin, who led reward modeling for Gemini at DeepMind, and Ioannis Antonoglou. Earlier this year it signed a compute deal with SpaceX for capacity at the Colossus 2 data center. The bet is legible: spend more compute in training so inference costs less, then sell Chinese-tier reasoning at Chinese-tier cost, built and hosted in the US.

Why the cost argument is the interesting one

Benchmark supremacy is a parade; margin is the business. Every point of reasoning quality Beam holds while burning a quarter of the compute is a point of margin for whoever hosts it. 2026 has been the year DeepSeek and Qwen undercut Western labs on price per token, and Reflection is trying to fill the one gap that leaves open: an American open model that competes on cost instead of capability.

The honest version of the headline: America has a new open model that costs less to run and scores like China's models from four months ago. Whether "costs less" beats "scores higher" is a question for whoever downloads the weights. They can't yet.

Sources

  1. [1] TechCrunch — “Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost” (Oct 5, 2026)Read source
  2. [2] Reuters via SRN News — “Nvidia-backed Reflection unveils first AI model to take on Chinese open models” (Oct 5, 2026)Read source
  3. [3] MarkTechPost — “Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads” (Oct 5, 2026)Read source