Model index
DeepSeek·DeepSeek V / R

DeepSeekMoE 16B

A 16.4B MoE using fine-grained expert segmentation and shared-expert isolation, released in Base and Chat variants.

LegacyLanguageCodingOpen weights
MODEL NUMBER023
Release date
2024-01-11
Organization
DeepSeek
Family
DeepSeek V / R
01

Specifications

Only published specifications are shown.

Total parameters
16.4B
Architecture
MoE
Context
4K
Input modalities
Text
Open weights
Yes
License
DeepSeek Model License
Status
Legacy
02

Overview

Overview

A 16.4B MoE using fine-grained expert segmentation and shared-expert isolation, released in Base and Chat variants.

Why it mattered

DeepSeekMoE was the key intermediate step validating DeepSeek’s efficient sparse-expert path before V2.

03

Architecture & capabilities

Key capabilities

moeEfficiencyMultilingual
04

Release history

Open weights

DeepSeekMoE 16B was released publicly, exploring efficient MoE through fine-grained and shared-expert routing.

View source
05

Related models

Select a card to view the related model.

No related models recorded yet

06

Sources