Reflection AI announced Beam on October 5: a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, text-only, built for coding, reasoning and agentic work. Reflection describes it as the lab’s first open-weight model and promises the weights under an Apache 2.0 license “later this month.” The weights are not out. As of October 5, the reflection organization on Hugging Face holds zero public models, no Beam license text has been published on any Reflection site we checked, and the technical report, model card and safety results are all listed as coming. What exists today is a blog post with benchmark tables, a waitlisted API beta and a coding-agent CLI released as prebuilt packages without source. Reflection AI
That puts Beam in a different category from Inkling, which Thinking Machines shipped in July with weights, a model card and Apache 2.0 terms on the same day. Every Beam spec and benchmark score here is Reflection’s own (the self-hosting figures further down are our estimates); no independent evaluation exists, and TechCrunch’s launch report says the performance claims “haven’t been independently verified.” The useful questions are what the table shows row by row, what the efficiency claim measures, and what a 501B checkpoint will take to run once it lands.
What is usable on October 5, and what is dated “later this month”
The blog post carries the specs: 52 layers, interleaved local and global attention and fine-grained routed experts, with the expert count and the number active per token undisclosed. Reflection AI Pretraining ran on 23.8 trillion tokens in what Reflection describes as “under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs”; reinforcement learning used 10.5K GB300 GPUs for four weeks and more than 100 million rollouts. The post says midtraining “extends Beam’s effective context length to 1M tokens,” but the API serves a 262,144-token context window, prompt and output combined, with output capped at 131,072, and the docs say that may change during the beta. Reflection developer docs Treat 1M as a training claim and 256K as the product.
The API is OpenAI-compatible, exposes the model ID Beam-501B-A23B, and keeps reasoning always on across five effort levels. Reflection developer docs Access is a waitlist; the beta is free with daily usage limits, and no per-token price appears on any public page we checked. Reflection beta pricing Mirror, Reflection’s coding agent, ships from a GitHub repository that “hosts beta releases and feedback, not source code,” and needs a Reflection API key to run Beam. reflection-oss/mirror-beta
Everything else is a date without a day: “We will release the weights, technical report, model card, and developer artifacts later this month.” Beam is “undergoing final red-teaming and evaluations,” and distribution partners are promised but unnamed. Reflection AI When we checked on October 5, OpenRouter’s model list had nothing from Reflection and Artificial Analysis had no Beam page.
Apache 2.0 is a promise, and the comparison set shows why the text matters
Reflection’s wording is “This month, we will release the weights under an Apache 2.0 license, along with documentation and the full stack for running, evaluating, and fine-tuning the model.” There is no weights repo, so no LICENSE file, and the Mirror repository reports no license either. Until a file exists, “promised Apache 2.0” is the most that can be printed.
The models in Reflection’s own table show how far a license file can drift from the headline. Moonshot’s Kimi K3 License requires a separate agreement from model-as-a-service operators with more than US$20M in revenue over twelve months; Qwen3.8-Max requires a separate license from model-as-a-service and AI work assistant businesses above US$50M; Z.ai’s GLM 5.3 license sends operators above US$10 billion through a security review, a clause we covered when GLM 5.3’s weights arrived about two weeks after the flagship reached the API. GLM 5.2 and DeepSeek V4.1 Flash are MIT, Inkling is Apache 2.0 and Nemotron 3 Ultra uses OpenMDW-1.1. Plain Apache 2.0 would put Beam in the permissive group; a modified one would not.
Promised dates have a mixed record. Moonshot kept its July 27 date for Kimi K3. Reflection’s own schedule has moved before: in October 2025 TechCrunch reported the company was “aiming to release” its first model “early next year,” and what arrived in October 2026 was an announcement.
Read the table row by row
Reflection’s tables compare Beam with seven models: Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. No eval settings are given: no reasoning effort, harness, pass@k or number of runs. The table gives no source for competitor scores. Only the efficiency figure’s caption says it drew other models’ evals from Artificial Analysis and DataCurve, and the AA Omniscience row is labelled “our results.” Many cells read NR, “scores that have not been reported,” so several rows compare Beam with only one to three other models rather than seven.
Where Beam leads: Inkling on most shared rows, with SWE Bench Pro v2-Hard at 77.2 versus 56.9 and Terminal Bench 2.1 at 80.1 versus 63.8, though it trails Inkling on IFBench (79.7 versus 79.8) and AA Omniscience (13.0 versus 14.2). It leads Nemotron 3 Ultra wherever both report except IFBench and an AA-LCR tie. Against GLM 5.2 the record splits: ahead on eight of thirteen shared rows, behind on five, including Terminal Bench (80.1 versus 81.0) and HLE without tools (36.2 versus 40.5).
Where it trails: Qwen 3.8 Max on every one of the thirteen shared rows, among them tau3 banking at 38.0 versus 55.2 and Terminal Bench at 80.1 versus 86.6. Kimi K3 by wide margins on BrowseComp and DeepSearchQA, both run with context management (77.4 versus 91.2 and 80.1 versus 95.0), and on DeepSWE (44.4 versus 68.0). DeepSeek V4.1 Flash on DeepSWE (74.2) and Terminal Bench (90.6). Reflection’s prose calls Beam “competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks” and concedes that “frontier open models like Kimi K3 remain ahead on raw capability.” Competitive with GLM 5.2 is a fair reading. “Approaching” Qwen means behind on every shared coding and agentic row, and on all thirteen shared rows overall.
The competitor cells need their own caution, because several do not match the vendors’ own model cards. The table lists GLM 5.2 at 44.0 on DeepSWE; Z.ai’s GLM 5.3 card lists it at 46.2, which would turn Beam’s 44.4 lead into a deficit. GLM 5.3 and Qwen are also listed below their own cards on that row. On MCP Atlas the table gives Kimi K3 82.3 and GLM 5.3 84.2, while Kimi’s card lists Kimi K3 at 84.2, so a column swap is possible. Kimi K3’s DeepSWE is 68.0 in the table and 67.5 on Kimi’s card; Inkling’s card has MCP Atlas at 74.1 against the table’s 76.0. GLM 5.3’s 61.0 and Kimi’s 68.0 also appear identically in the DeepSWE and SWE Atlas Codebase QnA rows, which looks like a copy error. The Terminal Bench cells match each vendor’s own card, though other labs’ cards list GLM 5.2 at 82.7 rather than 81.0. Read the table as Reflection-assembled, with undisclosed settings, not as a leaderboard. Z.ai GLM 5.3 model card
The efficiency claim is a FLOPs estimate over three benchmarks
Reflection’s strongest claim is that Beam “achieves scores comparable to GLM-5.2 while using 3–4× less inference compute” on advanced reasoning benchmarks. The figure’s caption defines the measure: FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt, excluding prompt prefill, context-dependent attention and serving overhead, “an approximate compute comparison rather than measured inference cost.” It covers DeepSWE, HLE and Terminal Bench 2.1 only.
Two things produce the ratio. GLM 5.2 runs about 40B active parameters to Beam’s 23B, roughly 1.7x by itself; the rest has to come from Beam generating fewer tokens per attempt, and the post describes a “controllable length penalty” in training that rewards success while discouraging unnecessary tokens. The reasoning effort used for the benchmarks is not disclosed. No price is published, so “lower cost” has no number behind it yet. Alex Heath reports Reflection saying Beam uses “roughly a quarter of GLM-5.3’s inference compute,” while the post names GLM-5.2; until Reflection says which model the comparison used, keep the two figures apart. Sources
What 501B means for self-hosting
No checkpoint exists to measure, so these are estimates: decimal gigabytes, weights only, the nominal 501B count, no KV cache or activations. At BF16 the weights are about 1,002 GB; at FP8 about 501 GB; at an ideal 4 bits about 250 GB. Real 4-bit checkpoints land closer to 4.9 to 5.1 bits per parameter once block scales are counted (the NVFP4 builds of Nemotron 3 Ultra and Inkling both land near that), which puts a practical 4-bit Beam around 305 to 320 GB. Per generated token the model reads roughly its 23B active parameters, about 46 GB at BF16 or 11.5 GB at 4 bits, a bandwidth figure rather than a capacity one. Our mixture-of-experts explainer covers the total-versus-active distinction.
Against published GPU memory figures, eight H200s (1,128 GB) hold BF16, eight H100s (640 GB) hold FP8, four H100s (320 GB) would hold a real 4-bit build’s weights with almost nothing left for cache, so they do not realistically fit, and a 512 GB unified-memory workstation holds 4-bit with roughly 200 GB left for cache and context. Two 128 GB machines do not fit. The KV cache size is unpublished, and no FP8 or NVFP4 release is promised, only “the full stack for running, evaluating, and fine-tuning the model.” The Open Models guide separates weight size from runtime memory and the machine guide covers headroom.
Where it sits among US open models
The shipped US field: Inkling at 975B total and 41B active, Apache 2.0, multimodal, public since July 15; Nemotron 3 Ultra at 550B and 55B active under OpenMDW-1.1, since June 3; gpt-oss, Apache 2.0, since August 2025, with no newer main model from OpenAI. Beam would be smaller than Inkling and Nemotron 3 Ultra in total parameters, with fewer active, and text-only where Inkling is not. Reflection’s table includes only Inkling and Nemotron 3 Ultra from that field, so “advances the Western open-weight frontier” is a claim about that pair.
The independent gap is larger than the table suggests. On Artificial Analysis’ current index (v4.3.2, rescaled since July, so not comparable with Inkling’s launch score of 41) Inkling sits at 25 against GLM-5.3 at 45 and Kimi K3 at 44. Beam has no score there. Reflection has the money to close it: $2 billion raised in October 2025 TechCrunch, October 2025, and a later round at a reported $25 billion pre-money valuation TechCrunch, October 2026. In a Davos interview in January, published by Alex Heath’s Sources on February 4, Heath wrote that Antonoglou said the startup plans to release what it hopes will be “the most powerful open-weight model in the world” later this year. By its own table, Beam is not that.
What to check when the weights land
- Hugging Face: a model repo under
reflection, whether it is gated, and whether the LICENSE file and the metadata both say unmodified Apache 2.0. config.json: total parameters against 501B, expert count and experts active per token, 52 layers, andmax_position_embeddings(256K or 1M).- Checkpoint format and size: about 1 TB at BF16; whether an FP8 or NVFP4 build appears, which the post does not promise.
- The technical report: effort, harness, pass@k and run counts per benchmark; whether the GLM 5.2 DeepSWE cell becomes 46.2; whether the duplicated pairs and MCP Atlas columns are corrected; the safety results.
- Independent scores: an Artificial Analysis entry, an OpenRouter listing, third-party hosts.
- Post-beta pricing and rate limits. The beta pricing page and the rate-limits doc disagree on the reset time, so plan around neither.
Until those items exist, Beam is a well-documented announcement from a lab that has not released an open-weight model before. Keep it off any evaluation shortlist until the repo appears, then run the checklist.
Key Details
| Spec | Detail |
|---|---|
| Lab | Reflection AI (founded 2024; Misha Laskin, CEO; Ioannis Antonoglou, CTO) |
| Announced | October 5, 2026 |
| Weights | Promised “later this month”; Hugging Face org had 0 public models on October 5 |
| License | Apache 2.0 promised; no license text published |
| Total Parameters | 501B (self-reported) |
| Active Parameters | 23B per token (about 4.6%) |
| Architecture | 52-layer sparse MoE; expert count and top-k undisclosed |
| Context | API window 262,144 tokens (prompt plus output), max output 131,072; 1M “effective” is a training claim |
| Modalities | Text only |
| Training | 23.8T pretraining tokens; RL on 10.5K GB300 GPUs, >100M rollouts (self-reported) |
| API | OpenAI-compatible, model ID Beam-501B-A23B, waitlist, free beta with daily limits, no public price |
| Self-Host (estimates) | ~1,002 GB BF16; ~501 GB FP8; ~305-320 GB practical 4-bit; KV cache unpublished |
| Independent Scores | None as of October 5 |
Sources
- Introducing Beam — Reflection AI
- Reflection organization overview — Hugging Face
- Developer platform docs: models, reasoning, rate limits — Reflection AI
- Beta pricing information — Reflection AI
- reflection-oss/mirror-beta — GitHub
- Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost — TechCrunch
- Reflection founders on the open-weight Beam release — Sources
- The new AI lab that’s coming for DeepSeek — Sources
- Reflection raises $2B to be America’s open frontier AI lab — TechCrunch
- GLM-5.3 model card — Z.ai
- Kimi K3 model card and license — Moonshot AI
- Qwen3.8-2.4T-A95B license — Alibaba Qwen
- Inkling model card — Thinking Machines Lab
- NVIDIA Nemotron 3 Ultra 550B-A55B — NVIDIA
- Artificial Analysis model pages: Inkling, GLM-5.3, Kimi K3 — Artificial Analysis
