Model index
Alibaba Cloud·Qwen

Qwen3.8-Flash-Next

A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.

CurrentReasoningCodingAgentMultimodalOpen weights
MODEL NUMBER076
Release date
2026-08-26
Organization
Alibaba Cloud
Family
Qwen
01

Specifications

Only published specifications are shown.

Total parameters
176B
Active parameters
6B
Architecture
Hybrid
Context
262K
Input modalities
Text · Image
Open weights
Yes
License
Apache-2.0
Status
Current
02

Overview

Overview

A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.

Why it mattered

It introduced QSA, Gated Residual, and N-gram memory to reduce compute while advancing long-context agents.

03

Changes from predecessor

Compared with Qwen3.8-Max

PreviousQwen3.8-Max2.4T · 1M ctx
CurrentQwen3.8-Flash-Next176B · 262K ctx
New capabilities Adjustable reasoning
04

Architecture & capabilities

Attention

Gated DeltaNetQwen Sparse Attention

Key capabilities

Adjustable reasoningCodingAgentLong context
05

Release history

Open weights

Qwen3.8-Flash-Next released open weights and became available as managed Qwen3.8-Flash.

View source
06

Related models

Select a card to view the related model.

07

Sources