Overview
A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.
A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.
Only published specifications are shown.
A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.
It introduced QSA, Gated Residual, and N-gram memory to reduce compute while advancing long-context agents.
Qwen3.8-Flash-Next released open weights and became available as managed Qwen3.8-Flash.
View sourceSelect a card to view the related model.