Overview
Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.
Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.
Only published specifications are shown.
Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.
DeepSeek-V2 showed that efficient sparse architecture could compete on both capability and cost in an open-weight model.
DeepSeek-V2 released model weights and its technical report.
View sourceSelect a card to view the related model.