Overview
A 2.8T open MoE using KDA and Attention Residuals with native vision and a 1M context window.
A 2.8T open MoE using KDA and Attention Residuals with native vision and a 1M context window.
Only published specifications are shown.
A 2.8T open MoE using KDA and Attention Residuals with native vision and a 1M context window.
K3 pushed open-model scale toward 3T parameters while unifying agent, vision, and long-context capabilities.
Kimi K3 was announced at 2.8T parameters with native vision.
View sourceSelect a card to view the related model.