Model index
DeepSeek·DeepSeek V / R

DeepSeek-V2

Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.

LegacyLanguageOpen weights
MODEL NUMBER016
Release date
2024-05-06
Organization
DeepSeek
Family
DeepSeek V / R
01

Specifications

Only published specifications are shown.

Total parameters
236B
Active parameters
21B
Architecture
MoE
Context
128K
Input modalities
Text
Open weights
Yes
License
DeepSeek License
Status
Legacy
02

Overview

Overview

Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.

Why it mattered

DeepSeek-V2 showed that efficient sparse architecture could compete on both capability and cost in an open-weight model.

03

Architecture & capabilities

Attention

MLA

Key capabilities

Long contextCoding
04

Release history

Open weights

DeepSeek-V2 released model weights and its technical report.

View source
05

Related models

Select a card to view the related model.

06

Sources