Overview
A hybrid MoE using Lightning Attention, with 456B total and 45.9B active parameters, trained at 1M context and extrapolatable to 4M.
A hybrid MoE using Lightning Attention, with 456B total and 45.9B active parameters, trained at 1M context and extrapolatable to 4M.
Only published specifications are shown.
A hybrid MoE using Lightning Attention, with 456B total and 45.9B active parameters, trained at 1M context and extrapolatable to 4M.
MiniMax-Text-01 released model weights.
View sourceSelect a card to view the related model.