Strata, a free, MIT-licensed inference engine and one-click installer, runs Qwen’s 125B-parameter Qwen3.8-Flash-Next on an ordinary gaming PC: one NVIDIA or AMD card with 12 GB of VRAM, 32–64 GB of system RAM and about 80 GB of SSD space. On the author’s RTX 5070 (12 GB) with 64 GB of RAM, the project reports 53–94 tokens per second of output depending on the quantization. The repo hit the Hacker News front page with roughly 635 points and ~296 comments, and it has passed 11,000 GitHub stars in under two weeks.
The speed numbers are the project’s own. 9to6AI has not hands-on tested Strata, and the fastest results use 2-bit and 3-bit quantizations of the model, which is where most of the debate sits.
Image credit: Strata
What shipped
- Product: Strata, an inference engine plus installer for Windows 10/11 and Linux (github.com/Niko1221/Strata), built by developer Niko1221 with community contributors. Latest release at the time of writing: v0.1.39.
- License: MIT for Strata itself. Each model file keeps its own license. Qwen3.8-Flash-Next ships under the Qwen Community License 1.0, not Apache or MIT.
- Model: Qwen3.8-Flash-Next from the Qwen team. It has 125B parameters with 6B active per token, plus a 51B n-gram embedding and a 4B multi-token-prediction (MTP) layer. Inside are 512 experts per layer (10 routed + 1 shared), 48 layers, 262,144 tokens of native context and support for up to 1M with YaRN. Qwen calls it an experimental preview of the architecture behind Qwen4.
- Interfaces: a local web app at
127.0.0.1:8080with Chat, Monitor and About tabs; an OpenAI-compatible API at/v1; Anthropic Messages at/v1/messages; and the OpenAI Responses API at/v1/responsesfor Codex CLI. Optional image input and an MCP server that lets coding assistants install, start and stop it. - Reasoning effort: none, low, medium or high, per request.
The headline numbers (project-reported)
Measured by the author on an RTX 5070 12 GB, Ryzen 5 7600 and 64 GB RAM. Output speed is for short chats, and prompt speed is for a 32K-token prompt.
- Q2_0: 94 tok/s output · 2,650 tok/s prompt
- IQ2_XS: 79 tok/s · 2,090 tok/s
- IQ3_XXS: 62 tok/s · 1,750 tok/s
- IQ3_S: 53 tok/s · 1,620 tok/s
- Coder (expert-pruned): 55 tok/s · 2,180 tok/s
On AMD (RX 9070 XT 16 GB, 47 GB RAM) the project lists 60 tok/s for Q2_0, 52 for IQ2_XS and 44 for the Coder build. Its detailed tables show Q2_0 still writing about 74 tok/s at 128K context and about 60 tok/s at the full 262K on the RTX 5070. The project estimates that an RTX 3090 (24 GB) should land around 100–140 tok/s, but that figure is a projection rather than a measurement.
On the Hacker News thread, the submitter reported about 124 tok/s on an RTX 4090 with 128 GB of DDR5. That is a single user report, not a controlled benchmark.
How it fits a 125B model on a 12 GB card
Strata leans on how mixture-of-experts models work. Only 10 of each layer’s 512 experts fire per token, so the full set never has to live in VRAM at once. Strata splits the model across the whole PC:
- GPU: attention and DeltaNet mixers, routers, shared experts, the output head, the MTP draft layer, the KV cache, plus an expert cache that fills spare VRAM with the most-used experts and adapts to the conversation. The docs say each extra GB of VRAM holds roughly 700 more experts.
- RAM: all 24,576 experts (512 × 48 layers). The CPU computes whatever the GPU cache misses, in parallel with the GPU.
- SSD: the model’s 28.8 GB n-gram embedding table. Only a few rows are read per token.
- Speculative decoding: the model’s own MTP layer drafts up to 3 tokens and the full model verifies them in one pass. The project reports 2.4–3.2 tokens per pass and 1.6–1.8× faster output with the same tokens.
- Long prompts: read in chunks of up to 8,192 tokens, with the next layer’s experts streamed over PCIe while the current layer runs.
None of these tricks is new on its own. The difference is that they are packaged for one specific model on consumer hardware. Strata borrows from llama.cpp/ggml and credits ideas from projects such as ninfer, Splash and HyperQwen.
Deployability: what you actually need
- GPU: NVIDIA RTX 20, 30, 40 or 50 series, or listed AMD Radeon cards (RX 7900 XT/XTX, 7800 XT/7700 XT, 9060 XT, 9070/9070 XT, Radeon AI PRO R9700, RX 6800/6900). 12 GB of VRAM is the stated target. The docs say 8 GB runs, but slowly. NVIDIA needs driver 580 or newer.
- RAM decides the model size: 32 GB means the Coder build. 48 GB fits IQ2_XS or Q2_0. 64 GB runs every standard size (IQ2_XS is recommended there). 96 GB or more leaves room for IQ3_S or Unsloth’s roughly 4-bit build.
- Disk: about 70–80 GB free. NVMe is strongly recommended.
- Not supported: macOS. Intel Arc and pre-RTX/older cards are only experimental community paths.
- First start: the PC can freeze for 1–3 minutes while 35–55 GB loads into RAM. The first message of a chat reads at about one minute per 30,000 tokens, and follow-ups start in seconds thanks to the conversation cache.
- Concurrency: one request at a time by default. Batching (
"parallel": N) is opt-in and slows each answer on a 12 GB card.
Picking a size: the trade-offs
Q2_0 and IQ2_XS are the speed picks. IQ3_XXS and IQ3_S are slower but are the quality picks in Strata’s own docs. All of them come from ISTA-DASLab’s GSQ-RCO quantizations of the original weights.
The Coder build keeps 256 of each layer’s 512 experts, chosen on code, agentic and vision calibration data, so it fits in 32 GB of RAM. Its authors report 91.3% of the full model’s SWE-bench Verified score and 98.7% of its LiveCodeBench v6 score. Strata’s README says it is weaker outside code, including on Chinese and other CJK text.
Swift 1.5 is a third-party fine-tune that thinks for much less time before answering. Its authors claim 63% fewer thinking tokens and answers 1.8× sooner with under 1% accuracy loss. It carries its own license, so check it before commercial use.
Qwen’s published scores for the full-precision model are strong, for example SWE-bench Pro 62.5, LiveCodeBench v6 91.9 and GPQA Diamond 91.7. A 2- or 3-bit build on your desktop is not the model Qwen benchmarked. Budget for some loss and measure it on your own tasks.
What it means for builders
Local coding agents get a serious option. Strata speaks the APIs your tools already use. For Claude Code, set ANTHROPIC_BASE_URL=http://127.0.0.1:8080. For Codex CLI, add a provider with wire_api = "responses". For OpenCode or any OpenAI-compatible app, point the base URL at http://127.0.0.1:8080/v1. Code, prompts and repos stay on your machine and there is no per-token bill.
RAM is the real budget line, not the GPU. A 12 GB card is enough to start. Going from 32 GB to 64 GB of system RAM is what unlocks the better sizes. If you are speccing a box for this, prioritize 64 GB of fast RAM and an NVMe SSD before a pricier card.
Benchmark before you switch workflows. The HN thread splits clearly. Some users report that the IQ3_XXS build beats their Qwen3.8-27B setups on coding. One commenter posted 114/128 versus 92/128 on a personal code benchmark. Others reject anything under 4-bit on principle. One user’s vision test showed much worse object-coordinate accuracy through Strata than through llama.cpp on the same weights, with a median error of 154.8 px against 46.5 px. These are single-user reports, so we have not verified them, but they point to the right test: run your own coding, tool-calling and image tasks at the quant you plan to use.
Mind the security defaults. Strata listens on 127.0.0.1 unless you change it, and v0.1.38 added protection against DNS-rebinding and cross-site requests when no API key is set. If you expose it on your LAN or through a tunnel, set an API key first. The README also suggests letting your AI coding assistant install it from the repo’s setup doc. That is convenient, but review what it runs, as you would with any install script.
Leave the “experimental speed projection” off unless you know what it does. Strata’s docs describe this optional control vector as a refusal-direction projection. It strips most refusals and removes a safety behaviour, and it also measurably shifts ordinary answers. It is off by default and should stay that way for work use.
What to watch
- Independent quality evals of the Q2/Q3 builds against full-precision Qwen3.8-Flash-Next on agentic coding and tool use.
- Whether llama.cpp or other mainstream runtimes absorb the same expert-caching approach. HN commenters note that llama.cpp’s contribution rules make heavily AI-written code hard to upstream.
- Vision accuracy through Strata’s image path versus llama.cpp’s multimodal stack.
- Support for the next Qwen generation, since Qwen positions Flash-Next as a preview of the Qwen4 architecture.
- Project maturity: release cadence is very fast (multiple versions in days), so pin a version for anything you depend on.
Key takeaways
- Strata (MIT) runs the 125B / 6B-active Qwen3.8-Flash-Next on a 12 GB NVIDIA or AMD GPU with 32–64 GB of RAM, on Windows or Linux.
- Project-reported speeds on an RTX 5070: 94 tok/s (Q2_0) down to 53 tok/s (IQ3_S). Those are low-bit quants, not the full-precision model.
- It works as a drop-in local backend for Claude Code, Codex CLI, OpenCode and OpenAI-compatible apps.
- System RAM, not VRAM, decides which quality tier you can run. 64 GB is the practical target.
- The model license is the Qwen Community License 1.0. Strata’s MIT license does not cover the weights.
- 9to6AI has not tested it. Validate quality on your own workload before replacing a hosted model.
Sources
- Strata — GitHub repository and README
- Strata — technical details (speeds, API, settings)
- Strata — how it works, credits and licenses
- Strata — releases
- Qwen3.8-Flash-Next — Hugging Face model card
- ISTA-DASLab GSQ-RCO quantizations — Hugging Face
- Run Qwen 3.8 Flash Next (125B) on consumer hardware — Hacker News



