ai insights

Reflection Announces Beam: 501B Open Weights Promised, Not Yet Released

Reflection announced Beam, a 501B/23B-active MoE with Apache 2.0 promised for later in October. No weights, license or report yet. What its tables show.

Reflection AI announced Beam on October 5: a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, text-only, built for coding, reasoning and agentic work. Reflection describes it as the lab’s first open-weight model and promises the weights under an Apache 2.0 license “later this month.” The weights are not out. As of October 5, the reflection organization on Hugging Face holds zero public models, no Beam license text has been published on any Reflection site we checked, and the technical report, model card and safety results are all listed as coming. What exists today is a blog post with benchmark tables, a waitlisted API beta and a coding-agent CLI released as prebuilt packages without source. Reflection AI

That puts Beam in a different category from Inkling, which Thinking Machines shipped in July with weights, a model card and Apache 2.0 terms on the same day. Every Beam spec and benchmark score here is Reflection’s own (the self-hosting figures further down are our estimates); no independent evaluation exists, and TechCrunch’s launch report says the performance claims “haven’t been independently verified.” The useful questions are what the table shows row by row, what the efficiency claim measures, and what a 501B checkpoint will take to run once it lands.

Announced versus available, October 5, 2026. Left, in solid green, what Reflection has published for Beam today: a blog post with specs and benchmark tables, a waitlist-only API beta that is free with daily limits, model ID Beam-501B-A23B serving 256K context, and the Mirror CLI as prebuilt packages that need a Reflection API key to use Beam; the same-day check found zero public models in Reflection's Hugging Face organization, no license file, no technical report and nothing on OpenRouter. Right, in dashed amber, what is promised for later in October: weights with the format unstated, an Apache 2.0 license with no text to read yet, the technical report and model card, safety results with red-teaming still running, and unnamed distribution partners. A bottom strip plots six of the models in Reflection's Terminal Bench 2.1 row: Inkling 63.8, Beam 80.1, GLM 5.2 81.0, Qwen 3.8 Max 86.6, Kimi K3 88.3, DeepSeek V4.1 Flash 90.6, as published in Reflection's table with no eval settings disclosed.

What is usable on October 5, and what is dated “later this month”

The blog post carries the specs: 52 layers, interleaved local and global attention and fine-grained routed experts, with the expert count and the number active per token undisclosed. Reflection AI Pretraining ran on 23.8 trillion tokens in what Reflection describes as “under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs”; reinforcement learning used 10.5K GB300 GPUs for four weeks and more than 100 million rollouts. The post says midtraining “extends Beam’s effective context length to 1M tokens,” but the API serves a 262,144-token context window, prompt and output combined, with output capped at 131,072, and the docs say that may change during the beta. Reflection developer docs Treat 1M as a training claim and 256K as the product.

The API is OpenAI-compatible, exposes the model ID Beam-501B-A23B, and keeps reasoning always on across five effort levels. Reflection developer docs Access is a waitlist; the beta is free with daily usage limits, and no per-token price appears on any public page we checked. Reflection beta pricing Mirror, Reflection’s coding agent, ships from a GitHub repository that “hosts beta releases and feedback, not source code,” and needs a Reflection API key to run Beam. reflection-oss/mirror-beta

Everything else is a date without a day: “We will release the weights, technical report, model card, and developer artifacts later this month.” Beam is “undergoing final red-teaming and evaluations,” and distribution partners are promised but unnamed. Reflection AI When we checked on October 5, OpenRouter’s model list had nothing from Reflection and Artificial Analysis had no Beam page.

Apache 2.0 is a promise, and the comparison set shows why the text matters

Reflection’s wording is “This month, we will release the weights under an Apache 2.0 license, along with documentation and the full stack for running, evaluating, and fine-tuning the model.” There is no weights repo, so no LICENSE file, and the Mirror repository reports no license either. Until a file exists, “promised Apache 2.0” is the most that can be printed.

The models in Reflection’s own table show how far a license file can drift from the headline. Moonshot’s Kimi K3 License requires a separate agreement from model-as-a-service operators with more than US$20M in revenue over twelve months; Qwen3.8-Max requires a separate license from model-as-a-service and AI work assistant businesses above US$50M; Z.ai’s GLM 5.3 license sends operators above US$10 billion through a security review, a clause we covered when GLM 5.3’s weights arrived about two weeks after the flagship reached the API. GLM 5.2 and DeepSeek V4.1 Flash are MIT, Inkling is Apache 2.0 and Nemotron 3 Ultra uses OpenMDW-1.1. Plain Apache 2.0 would put Beam in the permissive group; a modified one would not.

Promised dates have a mixed record. Moonshot kept its July 27 date for Kimi K3. Reflection’s own schedule has moved before: in October 2025 TechCrunch reported the company was “aiming to release” its first model “early next year,” and what arrived in October 2026 was an announcement.

Read the table row by row

Reflection’s tables compare Beam with seven models: Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. No eval settings are given: no reasoning effort, harness, pass@k or number of runs. The table gives no source for competitor scores. Only the efficiency figure’s caption says it drew other models’ evals from Artificial Analysis and DataCurve, and the AA Omniscience row is labelled “our results.” Many cells read NR, “scores that have not been reported,” so several rows compare Beam with only one to three other models rather than seven.

Where Beam leads: Inkling on most shared rows, with SWE Bench Pro v2-Hard at 77.2 versus 56.9 and Terminal Bench 2.1 at 80.1 versus 63.8, though it trails Inkling on IFBench (79.7 versus 79.8) and AA Omniscience (13.0 versus 14.2). It leads Nemotron 3 Ultra wherever both report except IFBench and an AA-LCR tie. Against GLM 5.2 the record splits: ahead on eight of thirteen shared rows, behind on five, including Terminal Bench (80.1 versus 81.0) and HLE without tools (36.2 versus 40.5).

Where it trails: Qwen 3.8 Max on every one of the thirteen shared rows, among them tau3 banking at 38.0 versus 55.2 and Terminal Bench at 80.1 versus 86.6. Kimi K3 by wide margins on BrowseComp and DeepSearchQA, both run with context management (77.4 versus 91.2 and 80.1 versus 95.0), and on DeepSWE (44.4 versus 68.0). DeepSeek V4.1 Flash on DeepSWE (74.2) and Terminal Bench (90.6). Reflection’s prose calls Beam “competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks” and concedes that “frontier open models like Kimi K3 remain ahead on raw capability.” Competitive with GLM 5.2 is a fair reading. “Approaching” Qwen means behind on every shared coding and agentic row, and on all thirteen shared rows overall.

The competitor cells need their own caution, because several do not match the vendors’ own model cards. The table lists GLM 5.2 at 44.0 on DeepSWE; Z.ai’s GLM 5.3 card lists it at 46.2, which would turn Beam’s 44.4 lead into a deficit. GLM 5.3 and Qwen are also listed below their own cards on that row. On MCP Atlas the table gives Kimi K3 82.3 and GLM 5.3 84.2, while Kimi’s card lists Kimi K3 at 84.2, so a column swap is possible. Kimi K3’s DeepSWE is 68.0 in the table and 67.5 on Kimi’s card; Inkling’s card has MCP Atlas at 74.1 against the table’s 76.0. GLM 5.3’s 61.0 and Kimi’s 68.0 also appear identically in the DeepSWE and SWE Atlas Codebase QnA rows, which looks like a copy error. The Terminal Bench cells match each vendor’s own card, though other labs’ cards list GLM 5.2 at 82.7 rather than 81.0. Read the table as Reflection-assembled, with undisclosed settings, not as a leaderboard. Z.ai GLM 5.3 model card

The efficiency claim is a FLOPs estimate over three benchmarks

Reflection’s strongest claim is that Beam “achieves scores comparable to GLM-5.2 while using 3–4× less inference compute” on advanced reasoning benchmarks. The figure’s caption defines the measure: FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt, excluding prompt prefill, context-dependent attention and serving overhead, “an approximate compute comparison rather than measured inference cost.” It covers DeepSWE, HLE and Terminal Bench 2.1 only.

Two things produce the ratio. GLM 5.2 runs about 40B active parameters to Beam’s 23B, roughly 1.7x by itself; the rest has to come from Beam generating fewer tokens per attempt, and the post describes a “controllable length penalty” in training that rewards success while discouraging unnecessary tokens. The reasoning effort used for the benchmarks is not disclosed. No price is published, so “lower cost” has no number behind it yet. Alex Heath reports Reflection saying Beam uses “roughly a quarter of GLM-5.3’s inference compute,” while the post names GLM-5.2; until Reflection says which model the comparison used, keep the two figures apart. Sources

What 501B means for self-hosting

No checkpoint exists to measure, so these are estimates: decimal gigabytes, weights only, the nominal 501B count, no KV cache or activations. At BF16 the weights are about 1,002 GB; at FP8 about 501 GB; at an ideal 4 bits about 250 GB. Real 4-bit checkpoints land closer to 4.9 to 5.1 bits per parameter once block scales are counted (the NVFP4 builds of Nemotron 3 Ultra and Inkling both land near that), which puts a practical 4-bit Beam around 305 to 320 GB. Per generated token the model reads roughly its 23B active parameters, about 46 GB at BF16 or 11.5 GB at 4 bits, a bandwidth figure rather than a capacity one. Our mixture-of-experts explainer covers the total-versus-active distinction.

Against published GPU memory figures, eight H200s (1,128 GB) hold BF16, eight H100s (640 GB) hold FP8, four H100s (320 GB) would hold a real 4-bit build’s weights with almost nothing left for cache, so they do not realistically fit, and a 512 GB unified-memory workstation holds 4-bit with roughly 200 GB left for cache and context. Two 128 GB machines do not fit. The KV cache size is unpublished, and no FP8 or NVFP4 release is promised, only “the full stack for running, evaluating, and fine-tuning the model.” The Open Models guide separates weight size from runtime memory and the machine guide covers headroom.

Where it sits among US open models

The shipped US field: Inkling at 975B total and 41B active, Apache 2.0, multimodal, public since July 15; Nemotron 3 Ultra at 550B and 55B active under OpenMDW-1.1, since June 3; gpt-oss, Apache 2.0, since August 2025, with no newer main model from OpenAI. Beam would be smaller than Inkling and Nemotron 3 Ultra in total parameters, with fewer active, and text-only where Inkling is not. Reflection’s table includes only Inkling and Nemotron 3 Ultra from that field, so “advances the Western open-weight frontier” is a claim about that pair.

The independent gap is larger than the table suggests. On Artificial Analysis’ current index (v4.3.2, rescaled since July, so not comparable with Inkling’s launch score of 41) Inkling sits at 25 against GLM-5.3 at 45 and Kimi K3 at 44. Beam has no score there. Reflection has the money to close it: $2 billion raised in October 2025 TechCrunch, October 2025, and a later round at a reported $25 billion pre-money valuation TechCrunch, October 2026. In a Davos interview in January, published by Alex Heath’s Sources on February 4, Heath wrote that Antonoglou said the startup plans to release what it hopes will be “the most powerful open-weight model in the world” later this year. By its own table, Beam is not that.

What to check when the weights land

  • Hugging Face: a model repo under reflection, whether it is gated, and whether the LICENSE file and the metadata both say unmodified Apache 2.0.
  • config.json: total parameters against 501B, expert count and experts active per token, 52 layers, and max_position_embeddings (256K or 1M).
  • Checkpoint format and size: about 1 TB at BF16; whether an FP8 or NVFP4 build appears, which the post does not promise.
  • The technical report: effort, harness, pass@k and run counts per benchmark; whether the GLM 5.2 DeepSWE cell becomes 46.2; whether the duplicated pairs and MCP Atlas columns are corrected; the safety results.
  • Independent scores: an Artificial Analysis entry, an OpenRouter listing, third-party hosts.
  • Post-beta pricing and rate limits. The beta pricing page and the rate-limits doc disagree on the reset time, so plan around neither.

Until those items exist, Beam is a well-documented announcement from a lab that has not released an open-weight model before. Keep it off any evaluation shortlist until the repo appears, then run the checklist.

Key Details

SpecDetail
LabReflection AI (founded 2024; Misha Laskin, CEO; Ioannis Antonoglou, CTO)
AnnouncedOctober 5, 2026
WeightsPromised “later this month”; Hugging Face org had 0 public models on October 5
LicenseApache 2.0 promised; no license text published
Total Parameters501B (self-reported)
Active Parameters23B per token (about 4.6%)
Architecture52-layer sparse MoE; expert count and top-k undisclosed
ContextAPI window 262,144 tokens (prompt plus output), max output 131,072; 1M “effective” is a training claim
ModalitiesText only
Training23.8T pretraining tokens; RL on 10.5K GB300 GPUs, >100M rollouts (self-reported)
APIOpenAI-compatible, model ID Beam-501B-A23B, waitlist, free beta with daily limits, no public price
Self-Host (estimates)~1,002 GB BF16; ~501 GB FP8; ~305-320 GB practical 4-bit; KV cache unpublished
Independent ScoresNone as of October 5

Sources

Continue reading.

Insight13 min read

Google Ships EmbeddingGemma 2: Text, Images, Video and Audio in One Vector Space

EmbeddingGemma 2 embeds text, code, image, video and audio in one 768-d space, ungated under Apache 2.0. Real sizes, self-reported scores, what to test.

Insight14 min read

Nano Banana 2.1: Half the Per-Image Price, 23 Days Left for Nano Banana 2 on the API

Google's Nano Banana 2.1 about halves the per-image price but triples input rates, and retires Nano Banana 2 on the Gemini API October 29. What to check.

Insight3 min read

Claude Sonnet 5.5: Compare Cost per Finished Task

Sonnet 5.5 keeps Sonnet 5 token prices. Its claimed savings come from using fewer tokens. How to test that claim in your workflow.