On August 26 Alibaba's Qwen team released the weights of Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model with only 6B active per token. It isn't a flagship release. It's the first open-weight model built on the architecture that will underpin Qwen4. Same playbook as Qwen3-Next before Qwen3.5: ship the architecture first, let the community take it apart, then build the full family on top of what works.
Two weeks in, the community has taken it apart. My read: the architecture matters more than the scores, and the scores need a date stamp before you read them.
What's new
Four changes, each aimed at a different scaling bottleneck. They sit inside a 48-layer layout arranged as 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE)), so only 12 of the 48 layers run full attention.
Qwen Sparse Attention (QSA). The hybrid design pairs Gated DeltaNet (linear attention) with full attention, same idea as Qwen3-Next, but the attention half is rebuilt. Standard sparse attention picks individual tokens to attend to; QSA picks micro-blocks of tokens instead, using a small indexer (4 query heads, 1 shared key head, budget of 2048 tokens). Block-level selection is cheaper to schedule on real hardware and plays better with paged KV caches. The payoff shows up at long context: Qwen reports up to 7.6x faster prefill and 4.9x faster decode at 1M tokens versus Qwen3.7-Plus, and up to 8.6x overall throughput with prefix caching. Long agentic sessions are exactly the workload that used to cost the most.
N-gram Embedding. Between the embedding parameters and the experts sits a 51B-parameter table of 20M n-grams (bigrams and trigrams, indexed at layer 2). It works like a giant lookup table for local patterns: the index is computed deterministically from the current token and its neighbors, so there is no learned routing and no gating matmul. Those parameters never enter the per-token compute budget, and because lookup is a memory access rather than a matrix multiply, they can live in system RAM instead of VRAM. vLLM ships a VLLM_PLE_CPU_OFFLOAD=1 flag that does exactly that. So the checkpoint carries roughly 180B parameters on disk once you count the table, and still behaves like a 6B-active model at inference.
Gated Residual. The residual stream, which every layer reads from and writes to, gets widened and fitted with gates: an element-wise, data-dependent read gate and a per-branch scalar write gate. It buys stability at depth with low inference overhead. Useful, but it mostly helps training rather than anything you feel at runtime.
Training recipe. Muon and AdamW applied to different weight categories, and no batch-size warmup at all. Training starts directly at target batch size, guided by refitted scaling laws. DataCamp estimates the model cost about a ninth of what training Qwen3.7-Plus took.
The numbers
Straight from the model card, so self-reported and bolded by Qwen:
| Benchmark | Qwen3.8-Flash-Next | Claude Opus 4.6 Max | DeepSeek V4 Flash 0731 |
|---|---|---|---|
| SWE-bench Pro | 62.5 | 53.4 | 56.0 |
| DeepSWE 1.1 | 58.7 | -- | 54.4 |
| CoWorkBench | 73.9 | 68.2 | 45.1 |
| JobBench | 55.7 | 36.6 | 41.3 |
| GPQA Diamond | 91.7 | 91.3 | 90.8 |
| HLE | 35.9 | 40.0 | 33.8 |
Claude Opus 4.6 Max is the February model, so the Claude column is last season's frontier, not this one. That is the same pattern I saw in DeepSeek V4 Flash 0731: a model you run on your own hardware is now scoring above what the frontier was shipping a few months earlier. On the agentic rows the gap is clear, 62.5 against 53.4 on SWE-bench Pro and 19 points on JobBench (55.7 vs 36.6), the widest margin in the table. The exam rows still favor the closed model, HLE 35.9 vs 40.0, and on Agents' Last Exam pass@1 even V4 Flash 0731 edges Qwen out (25.2 vs 24.3).
It's also multimodal, and the vision numbers look just as one-sided as the coding ones: AndroidWorld 84.5 (Opus: 62.0), RealWorldQA 88.5 (Opus: 73.9), MathVision up to 95.7 with code interpreter (Opus: 65.5).
Running it locally
I'm running it on DS4 as I write this, via PR #991 (not mine; @ivanfioravanti's PR adds Metal inference, external PLE weights, and optional MTP decoding, and it is still open as of this writing). It beats V4 Flash 0731 on my machine on both axes: on the mixed q2/q4 quant, V4 Flash gives me 220 tokens/s prefill and 22 tokens/s generation, while Qwen3.8-Flash-Next does 450 tokens/s prefill and 37 tokens/s generation. Twice the prefill and 1.7x on generation, on the same machine and the same engine. Caveat: those are best-case numbers on a nearly empty context window. Long prompts are where the architectures should separate (QSA is built for exactly that), but I haven't measured it yet.
That reverses the DeepSeek post, where I wrote that DSpark did nothing for my generation speed no matter how I tuned the parameters. Qwen's answer to the same bottleneck, a multi-token prediction head trained into the model itself rather than bolted on beside it, does make it faster on the same hardware.
The architecture explains why. Only a quarter of the layers run full attention, the n-gram embedding table is RAM-resident, and 6B active parameters per token is less than half of what V4 Flash activates. QSA's block-level selection is also kinder to the scheduler than token-level sparse attention; long-context prefill is where all of that compounds.
More broadly, Ollama ships a qwen3.8-flash-next:125b-mlx build at 105 GB with a 262K context window, which fits a 128 GB Apple Silicon machine. On the AMD side, r/LocalLLaMA reports 120 tokens/s generation and 12,000 tokens/s prefill on a 4×R9700 rig with optimized vLLM, single request. Unsloth has quantizations for smaller machines, and the n-gram table offloads to CPU RAM on vLLM, so those 51B "extra" parameters don't compete for your GPU memory.
MoE plus a tiny active count plus offloadable embeddings favors unified-memory machines and mixed CPU/GPU rigs. The compute per token is modest; only the memory footprint is large. In a year when RAM is absurdly expensive, the models that need less compute per token are the ones that stay cheap to run once the memory is already paid for.
What it costs
On Qwen Cloud the hosted version (qwen3.8-flash) runs $0.16 per million input tokens and $0.47 output. Its direct rival GLM-5.3-Flash, released the same day as the Qwen weights, is cheaper on input during its promo window and has 1M context natively, but it activates 18B parameters per token against Qwen's 6B. Self-hosting is where that difference pays, and self-hosting is what this blog is about.
The latency gap is a hardware problem. My box is a 128 GB Mac Studio that cost about $4k, which counts as cheap in this RAM market, and it runs the model at 37 tokens/s. The 120 tokens/s figure above came off a 4×R9700 rig, and hardware in that bracket costs tens of thousands of dollars. So $4k gets you a model that works, and the version that feels like a hosted endpoint costs a lot more. 37 tokens/s is workable but slow enough that I notice it, and if the box is not already sitting on your desk, the hosted route stays the cheaper answer.
My take
Seven months is a long time at the frontier, and beating last season's flagship tells me nothing about what the labs are running today. The claim I'd actually defend is a smaller one. A model with 6B active parameters did real agentic coding work on hardware I own, and two weeks into my own projects it is still doing it, with latency rather than capability as the price.
I keep coming back to the n-gram embedding. Roughly 51B of the 180B parameters on disk sit in a table that does no matrix multiplication and can live in RAM. Whatever Qwen4 turns out to be, I'd bet it scales that axis further, because it's the one axis that doesn't fight the two things limiting local AI in 2026: memory prices and compute per token.
Local coding models are good enough to use now, and that matters more to me than the benchmark table. If you already own a 128 GB machine, download the MLX build. If you don't, run qwen3.8-flash through the API and revisit when the next preview lands.