DeepSeek-V2
DeepSeek-V2 showed that efficient sparse architecture could compete on both capability and cost in an open-weight model.
FAMILY / DEEPSEEK
A model family centered on efficient MoE scaling, open weights, and reasoning-focused post-training.
01 / LINEAGE MAP
02 / GENERATIONS
DeepSeek-V2 showed that efficient sparse architecture could compete on both capability and cost in an open-weight model.
V3 pushed open MoE general and coding capabilities toward the frontier and became the base for subsequent DeepSeek reasoning work.
R1 brought open-weight reasoning models into mainstream global comparison and accelerated replication of RL-based reasoning recipes.
A general model unifying non-thinking and thinking modes with sparse attention and everyday agent capability.
A high-reasoning V3.2 variant for math, competitive programming, and longer deliberation.
It introduced V4’s sparse-attention, million-token architecture and released the weights alongside the preview.
The efficient V4 preview with 284B total, 13B active parameters, and a 1M context window.
The production API version, re-post-trained on the same architecture as V4-Flash-Preview with stronger agent capability.
The production V4-Pro snapshot with stronger agents, low/high/max reasoning effort, and Responses API support.
An experimental V4 Flash API model adding image input.