AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.