Image credit: Reflection

Reflection Unveils Beam, a 501B Open-Weight Model Aimed at Coding Agents

Reflection AI has announced Beam, its first open-weight model. It is a sparse mixture-of-experts (MoE) model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic work. Reflection says it will release the weights under Apache 2.0 “later this month,” along with a technical report, model card and tooling to run, evaluate and fine-tune it. For now, Beam is only available through a waitlisted early-access API.

The Reflection Beam pitch is not “best open model.” Reflection’s own table puts Beam behind Kimi K3, Qwen 3.8-Max and DeepSeek V4.1 Flash on several tests. The claim is efficiency: GLM-5.2-level reasoning for 3–4× less inference compute, from a US lab. Every number below is Reflection’s own. Nobody outside the company has verified them yet, and 9to6AI has not tested Beam.

Image credit: Reflection

Key points

  • Beam: 501B total / 23B active MoE, text-only, pretrained on 23.8 trillion tokens. The API lists a 256K context window during beta and a June 2026 knowledge cutoff.
  • Weights are promised under Apache 2.0 later this month, with no exact day given. Today there is a waitlist and an OpenAI-compatible beta API.
  • Reflection’s benchmarks are mixed. Beam scores 80.9 on SWE-bench Verified and 80.1 on Terminal-Bench 2.1, against 81.0 for GLM 5.2 and 90.6 for DeepSeek V4.1 Flash on the latter.
  • The headline claim is compute efficiency, estimated from active parameters × generated tokens. It is not a measured cost per task.
  • The announcement reached the Hacker News front page with roughly 336 points and ~97 comments. Much of the thread was skeptical of the benchmark framing.

What shipped

  • Model: Beam, model ID Beam-501B-A23B. It is a sparse MoE with fine-grained routed experts and interleaved local and global attention across 52 layers. It works with text only, and Reflection notes it can still handle other modalities when they are converted to text.
  • Training scale (company-reported): pretraining on 23.8 trillion tokens from the web, public sources and “proprietary licensed datasets,” run in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Then came a reinforcement-learning (RL) run on 10,500 GB300 GPUs for four weeks. That run produced more than 100 million rollouts, used about 1.3 billion sandboxes and drew on nearly one million coding, agentic and STEM environments.
  • Context: the blog says midtraining extended effective context to 1M tokens. The developer docs list 256K with up to 128K output tokens for the beta API, and note that the limit may change.
  • Reasoning control: Beam always reasons. In the API, reasoning_effort takes low, medium (default), high, xhigh or max.
  • API (beta, waitlist): an OpenAI-compatible endpoint at https://api.reflection.ai/openai/v1 that supports Chat Completions and Models only, with streaming, tool calling and structured outputs. The Responses, Embeddings, Batch and Files endpoints are not supported. No public pricing has been published.
  • Coding agents: Reflection’s docs include setup guides for its own Mirror CLI and for Pi, OpenCode and Hermes through the OpenAI-compatible provider.
  • License: Apache 2.0 is promised for the weights. Reflection has not said it will release the training data.

The benchmarks, read honestly

These are Reflection’s published numbers. “NR” means no score was reported. Reflection drew the comparison scores from Artificial Analysis and DataCurve.

Benchmark Beam GLM 5.2 Kimi K3 Qwen 3.8-Max DeepSeek V4.1 Flash
Terminal-Bench 2.1 80.1 81.0 88.3 86.6 90.6
SWE-Bench Pro v1 65.5 62.1 NR 67.7 NR
DeepSWE v1.1 44.4 44.0 68.0 51.0 74.2
MCP Atlas 78.7 77.8 82.3 84.5 NR
tau3 banking 38.0 37.1 37.1 55.2 NR
HLE (no tools) 36.2 40.5 46.9 43.6 39.1
GPQA Diamond 90.5 91.2 93.5 92.6 90.9

Beam also lists 80.9 on SWE-bench Verified, where the only other reported scores are Thinking Machines’ Inkling at 77.6 and NVIDIA’s Nemotron 3 Ultra at 70.7. On BrowseComp with context management it scores 77.4, against 77.1 for Inkling and 91.2 for Kimi K3.

The pattern is clear. Beam roughly ties GLM 5.2, beats the other Western open models in the table (Inkling and Nemotron 3 Ultra) on most coding and reasoning rows, and trails the strongest Chinese open models on raw scores. Reflection says so directly: “Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time.” HN commenters pointed out that the headline charts compare mostly against weaker models, while GLM 5.3 and DeepSeek V4.1 Flash appear only in the full table.

What the efficiency claim actually measures

Reflection says Beam matches GLM-5.2 on advanced reasoning benchmarks while using 3–4× less inference compute, and that the gap is wider against 2T-plus-parameter models such as Qwen 3.8-Max. The method matters. Reflection estimates compute as 2 × active parameters × mean generated tokens per attempt, counting both reasoning and answer tokens. Its own footnote says this leaves out prompt prefill, context-dependent attention and serving overhead.

So the claim combines two real things: fewer active parameters (23B against roughly 40B for GLM 5.2) and shorter reasoning traces. Reflection trained Beam with a controllable length penalty and says the model learned to solve tasks with fewer tokens. Still, it is a model-compute estimate. It does not measure what you will pay per task or how fast you will get answers on your own hardware. That depends on pricing that hasn’t been published and on runtimes that don’t support Beam yet.

How Reflection says it built it

  • RL at scale: fully asynchronous policy-gradient training that Reflection says stays stable even on rollouts generated more than a day earlier (107 weight versions behind the current policy). It averaged about 110,000 concurrent rollouts and peaked at 170,000 concurrent sandboxes.
  • Environment curation: roughly one million tasks, mostly synthetic, filtered so they are neither trivially solvable nor impossible. Reflection says weaker data quality caused capability plateaus, and that the RL run ended “with no sign of saturation.”
  • Pretraining stability: auxiliary-loss-free load balancing (building on DeepSeek’s approach) with near-uniform expert use, SandwichNorm, FP32 residual accumulation and 92.3% goodput late in the run.
  • Safety: a separate safety-and-alignment “teacher” model, merged with the RL teacher through multi-teacher on-policy distillation, plus deliberative-alignment training. Reflection says it will publish safety eval results in the technical report and open-source its internal safety evals.

Who Reflection is

Reflection AI was founded by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. TechCrunch reports that it has raised about $4.7 billion from backers including NVIDIA, Sequoia and Lightspeed, citing PitchBook. TechCrunch also reports more than $7 billion in GB300 compute deals with SpaceX and Nebius. The company pitches Beam as a “Western open-weight” option and wants enterprises and governments to train custom “AI factory” versions on their own data. It is a different company from the team behind the disputed “Reflection 70B” release, a mix-up that came up again in the HN thread.

What it means for builders

You can’t download it yet. Until the weights land, Beam is a waitlisted beta API with unpublished pricing and limits that may change. Don’t plan a migration around it. Plan an evaluation.

This is a server model, not a desktop model. Our own back-of-envelope math, for weights only and before KV cache: 501B parameters come to roughly 1 TB at 16-bit, around 500 GB at 8-bit and about 250–300 GB at 4-bit. Self-hosting means a multi-GPU node or heavy quantization. Only 23B active parameters per token helps throughput, but all the experts still need to sit somewhere.

The real draw is licensing and provenance. If your company or client cannot use Chinese-origin weights for procurement or policy reasons, a permissively licensed US model around GLM 5.2’s level is a meaningful new option. If you have no such constraint, the published numbers suggest DeepSeek V4.1 Flash, Kimi K3 or Qwen 3.8-Max are still the stronger open picks on raw capability.

Low-friction to test. Because the endpoint speaks OpenAI Chat Completions, most agent harnesses only need a new base URL, API key and model ID. Tools that depend on the Responses API won’t work against it as documented.

Measure tokens, not just scores. If you get access, log reasoning plus output tokens per solved task at each effort level, next to your current model. That is the efficiency claim you can verify yourself.

What to watch

  • The actual weight drop on Hugging Face: the license file, any usage restrictions and the release of the technical report with safety evals.
  • Independent evals from Artificial Analysis, LMArena and agentic-coding leaderboards, especially Terminal-Bench and SWE-Bench Pro, run by third parties.
  • Runtime support in vLLM, SGLang and llama.cpp, plus community quantizations. These will decide whether Beam is practical to self-host.
  • API pricing and which hyperscalers or neoclouds carry it at launch.
  • Reflection says the next model in the series is already training.

Sources

Leave a Comment

Your email address will not be published. Required fields are marked *