02 / MODEL INDEX
AI model index
Browse first release dates, organizations, families, architecture, context length, and weight availability for every model in the archive.
The first natively multimodal GLM-5 Flash model: 320B total, 18B active parameters with hybrid sparse and linear attention.
A Qwen4 architecture preview: 125B main model plus 51B N-gram embeddings, 6B active per token, extendable to 1M context.
An experimental V4 Flash API model adding image input.
Kept the GLM-5.2 base and improved complex coding, long-horizon tasks, and cyber capability entirely through scaled post-training.
The 27B dense open Qwen3.8 model for local deployment, vision, coding, and tool use.
The current Gemini Flash workhorse for software engineering, web development, multi-step agents, and design fidelity.
The production V4-Pro snapshot with stronger agents, low/high/max reasoning effort, and Responses API support.
Built on Grok 4.5 with stronger long-running agents, complex-codebase work, and interactive visual projects.
A 30B open model for always-on local agents, designed to run on one consumer GPU or a Mac or PC.
The current preview all-in-one video model, generating up to 30 seconds with as many as 20 multimodal references.
The current Muse flagship, with stronger coding and the multi-agent Muse Code tool.
The Qwen3.8 flagship with 2.4T total, 95B active parameters, text-image-video input, and 1M context.
The production API version, re-post-trained on the same architecture as V4-Flash-Preview with stronger agent capability.
Generates up to 30 seconds with multi-round extension, 30 image/10 video/10 audio references, and timestamp editing.
An open omni-modal generation system producing up to 15-second, 2K, 24fps video with native stereo audio.
The new Opus flagship with a default 1M context and adaptive thinking for long-running agents, research, and enterprise work.
The current default, focused on aesthetics, image quality, personalization, and fewer low-quality results.
A low-latency 3.5 tier designed for high-volume automation and subagents.
Improved token efficiency, coding, and agent planning over 3.5 Flash while reducing output verbosity.
Targets directly usable visual content with 4.5K-token prompts, 10px text, multilingual rendering, and complex layouts.
A 2.8T open MoE using KDA and Attention Residuals with native vision and a 1M context window.
A fast flagship for coding, agents, and knowledge work, available in Grok Build and the API.
Improved tools, computer use, coding, and multimodal understanding with a 1M-context Meta Model API preview.
The professional-design flagship for information visualization, pixel-level interactive editing, realism, and multilingual generation.
An agentic image model that uses search and code, self-refines, and supports single- and multi-image references.
A preview Muse video model emphasizing fidelity, temporal consistency, and native audio; not yet generally available.
The final Hy3 used product feedback to improve agents, coding, and output reliability.
A new Sonnet generation with a default 1M context, adaptive thinking, and stronger agent capability closer to Opus.
The fastest and lowest-cost high-throughput model in the Nano Banana family.
Google’s Gemini-native model for cost-efficient video generation and conversational editing.
A three-tier series—Sol, Terra, and Luna—covering flagship, balanced, and cost-efficient capability levels.
The current Seed line, improving long-running general agents, end-to-end coding, and scientific workflows.
Extended context to 1M and optimized for large codebases, long documents, and project-scale agent work.
An experimental Gemma 4 text-diffusion model exploring lower latency through parallel block generation.
A Mythos-class model made generally available with conservative safeguards for software engineering, research, and complex knowledge work.
The current Ray model for fuller creative control and production API workflows.
A unified encoder-free multimodal model for laptops, filling the gap between E4B and 26B.
MiniMax’s current open flagship with 1M context, native image/video understanding, coding, and computer use.
Improved agent judgment, tool efficiency, computer use, and long-session collaboration over Opus 4.7.
The first public model in Google’s 3.5 generation, combining frontier intelligence with low-latency action in the Flash tier.
The low-latency, high-throughput, cost-efficient tier in the Gemini 3 family.
The everyday fast tier of GPT-5.5, improving accuracy, concision, and use of personalized context.
The high-capability DeepSeek V4 preview: a 1.6T-total, 49B-active MoE with a 1M context window.
The efficient V4 preview with 284B total, 13B active parameters, and a 1M context window.
The first Hy3 preview after Tencent rebuilt its training stack, combining fast/slow reasoning with sparse MoE.
A 1M-context frontier model for complex knowledge work, computer use, coding, and scientific research.
MiMo’s current general model combining native full-modal understanding, 1M context, and agent execution.
The high-intensity MiMo-V2.5 agent tier for long-range reasoning, coding, and complex-task efficiency.
The 27B dense Qwen3.6 tier for local deployment and general multimodal agents.
The current flagship for 2K production assets, complex layouts, multilingual text, and high-fidelity editing.
The open successor to K2.5 with stronger long-horizon coding, general agents, vision, and Agent Swarm.
Focused on hard software engineering, long-run consistency, higher-resolution vision, and self-verification.
The efficient Qwen3.6 MoE tier, using 3B active parameters for multimodal, coding, and tool tasks.
Added native 2K HD, roughly 4–5× faster standard generation, and stronger text and detail retention.
The first Muse model, with native multimodal reasoning, tools, visual chain of thought, and multi-agent orchestration.
A GLM-5 update for long-horizon work, designed to execute autonomously for hours within one task.
Unified audio text-to-video, first/last frames, continuation, references, and editing in a hosted 1080p line.
The high-quality dense Gemma 4 flagship for local reasoning, coding, and agent workflows.
The Gemma 4 MoE activating only 3.8B parameters per step for local speed and intelligence density.
Wan’s image model for image sequences, multi-reference, interactive editing, and output up to 4K.
Strengthened agent teams, self-evolution, and end-to-end software productivity.
A trillion-parameter, 42B-active agent flagship with a 1M context window.
Unified text, vision, and speech perception with tools and GUI control in one agent foundation.
A fast small model for coding and subagents with text, image, tools, and a 400K context window.
The smallest GPT-5.4 tier for high-throughput extraction, ranking, and simple coding subtasks.
Unified GPT-5.3-Codex coding advances with general reasoning, tools, and professional document work.
A GPT-5.3 update for everyday conversation, focused on fewer unnecessary refusals, better search accuracy, and smoother writing.
Brings 0.5K–4K generation, thinking, and image-search grounding at Flash speed and price.
The successor preview to Gemini 3 Pro, with a separate endpoint tuned to prioritize custom tools.
A full Sonnet 4.5 upgrade across coding, computer use, long context, agent planning, and design.
The first Qwen3.5 flagship, extending the Qwen3-Next hybrid architecture into a native multimodal agent model.
Systematically improved multimodal understanding, foundation reasoning, and production agent performance.
Brought deeper reasoning and live search into lightweight image generation and editing.
Scaled to 744B total parameters with DSA, repositioning the model from code generation toward complex systems engineering.
A small, ultra-fast Codex research preview designed for real-time interactive coding.
A high-efficiency update for coding, tools, and workplace productivity.
Unified text, image, audio, and video references and editing with 15-second multi-shot stereo output.
Unified generation and editing with 1K-token complex prompts, native 2K, and more realistic detail.
A Codex model extending agentic coding into research, tool use, and end-to-end computer work.
Brought a 1M context beta to Opus with stronger long-running agents, code review, and adaptive thinking.
The native multimodal 3.0 line unifies image, video, and audio understanding, generation, and editing in narratives up to 15 seconds.
A natively multimodal continuation of K2 with visual coding and an Agent Swarm of up to 100 subagents.
Brought Ray3 to native 1080p while substantially reducing latency and cost.
A sub-second generation and editing model for consumer GPUs.
Improved real-world coding, tool use, and agent reliability.
The open successor to GLM-4.5 with continued gains in coding, reasoning, and agent tasks.
A GPT-5.2 variant optimized for long-horizon software engineering, large code changes, Windows, and context compaction.
A general agent model for search, code, GUI interaction, and complex workflows.
The low-latency Gemini 3 preview with stronger visual-spatial reasoning and agentic coding.
Used hybrid sliding-window attention and MTP for faster long-context coding and agents.
Improved precise editing, detail preservation, typography, and generation speed by roughly four times.
Added joint audio-video generation, multilingual dialogue, camera control, and fuller narrative expression.
A GPT-5 series upgrade for professional work and long-running agents, offered in Instant, Thinking, and Pro tiers.
A general model unifying non-thinking and thinking modes with sparse attention and everyday agent capability.
A high-reasoning V3.2 variant for math, competitive programming, and longer deliberation.
Runway’s current flagship for cinematic fidelity, complex sequenced prompts, and professional HDR output.
Unified video generation, references, and editing across Kling O1/2.6 with native audio.
The second-generation visual family with multi-reference editing, precise color control, and stronger realism.
Introduced effort control with stronger coding, agents, computer use, and multi-agent coordination.
A professional Gemini 3 Pro image model with reasoning, real-world knowledge, multilingual text, and 4K control.
A 2M-context fast agent model for real-world tool use and deep research, offered in reasoning and non-reasoning variants.
The first Gemini 3 Pro preview for complex reasoning, coding, tools, and multimodal work.
Used large-scale reinforcement learning to improve style, intent understanding, collaboration, and factuality in everyday conversation.
An adaptive-reasoning API update that reduced latency on easy tasks and added apply_patch and shell tools.
A K2-based thinking agent trained to reason while using tools natively.
Improved body motion, micro-expressions, stylization, and motion prompts with a lower-cost Fast variant.
Shifted MiniMax’s main line toward efficient coding and agent execution.
A low-latency Haiku model offering near-Sonnet-4 coding and computer-use capability at lower cost.
Improved narrative control, reference consistency, and editing, adding 1080p, 4K, and vertical output.
Improved physics, control, and multi-shot state persistence with native dialogue and sound; the product ended on 2026-04-26.
Strengthened long-running coding, computer use, and complex agents alongside the Claude Agent SDK.
A native multimodal MoE image model with 80B total and 13B active parameters.
A trillion-plus-parameter Qwen3 flagship for coding, agents, and million-token professional work.
An efficient Grok 4 variant unifying reasoning and non-reasoning modes with 2M context and live search.
Added visual reasoning, self-evaluation, native HDR EXR, and draft-to-4K workflows.
An 80B-total, 3B-active architecture preview combining hybrid linear attention and ultra-sparse MoE.
Unified text-to-image and editing, using a vision-language model for world knowledge and supporting up to 4K.
Applied Gemini 2.5 Flash knowledge and conversation to fast image generation and editing.
GPT-5 unified fast responses, deeper reasoning, and automatic routing, with flagship, mini, and nano variants in the API.
The lower-cost GPT-5 tier for workloads balancing reasoning quality, throughput, and price.
The smallest GPT-5 tier for classification, extraction, ranking, and lightweight subtasks.
A 20B MMDiT image foundation model focused on complex Chinese/English text rendering and precise editing.
A 355B-total, 32B-active hybrid-reasoning MoE unifying reasoning, coding, and agent capabilities.
Introduced dual-expert MoE video diffusion with open 720p text, image, speech, and animation variants.
A 1T-total, 32B-active open MoE focused on agentic tool use and coding.
Scaled reinforcement learning beyond Grok 3 Reasoning with native tools, web search, and X search.
Used MatFormer and Per-Layer Embeddings to bring multimodal understanding to mobile devices.
Introduced adaptive reasoning on a 230B MoE foundation unifying text and vision.
Animates one image into five seconds, extendable to 21 seconds with end frames, loops, and motion modes.
Advanced Hailuo with native 1080p, complex motion, and stronger physics.
An open reasoning model with 1M context using hybrid linear and softmax attention.
Established Seedance with native multi-shot storytelling, 1080p, and stable complex motion.
Unified generation and contextual editing with multi-turn changes, character consistency, and local control.
The Claude 4 flagship, focused on long-running coding, complex agents, and sustained multi-hour execution.
The balanced Claude 4 tier with stronger coding, reasoning, instruction following, and parallel tool use.
Supports up to 2K and multiple aspect ratios with stronger detail, spelling, and typography.
Added native dialogue, ambience, and sound effects to Veo.
Xiaomi’s first open reasoning model, exploring the full pretraining-to-RL path at 7B scale.
The flagship Qwen3 MoE, combining thinking and non-thinking modes in one open model.
Brought GPT-4o-native image generation to the API with world knowledge, style adherence, and typography.
A natively multimodal MoE with 17B active parameters, 16 experts, and up to 10M context.
The higher-capability Llama 4 open model with 400B total parameters and 128 experts.
Added default personalization, Omni Reference, 10× draft mode, and conversational creation.
Improved subject, location, and style consistency across scenes with more direct camera control.
Google built “thinking” directly into the general Gemini generation, with emphasis on complex reasoning and coding.
Added image input, a 128K context window, and support for more than 140 languages.
Released open 1.3B and 14B text/image-to-video weights plus VACE general video editing.
Combined near-instant responses and controllable extended thinking in one model, alongside Claude Code.
xAI’s reasoning-agent generation, extending complex-task capability through Think and DeepSearch modes.
A low-latency multimodal model that introduced native tool use and the Flash Thinking experimental path.
Used large-scale reinforcement learning to shape reasoning and released both full and distilled model weights.
Moonshot’s multimodal reasoning work exploring long-context reinforcement learning and test-time reasoning scale.
Luma’s second-generation large video model with stronger natural motion, physics, and camera expression.
A 671B-total, 37B-active MoE model that extended the efficient-training and open-release path.
Advanced Google’s video line with stronger physics, cinematic language, and research output up to 4K.
Sora’s first production release with up to 20 seconds, 1080p, extension, blending, and storyboards.
Brought near-Llama-3.1-405B instruction quality to a much smaller 70B model.
Tencent’s first large open MoE with 389B total and 52B active parameters.
A customizable open image tier spanning 8.1B Large, Turbo, and 2.5B Medium variants.
Added 11B and 90B vision models plus 1B and 3B edge-oriented variants.
A major Google text-to-image generation with better detail, lighting, and prompt adherence.
BFL’s first flow-matching image family across open developer, fast, and hosted professional variants.
Llama’s first 405B open flagship, extending the family to 128K context.
Raised reasoning quality in 9B and 27B sizes designed for accessible accelerators.
Introduced MMDiT in a 2B model aimed at consumer GPUs.
The largest dense Qwen2 model, using GQA across the family and extending support to 27 additional languages.
An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.
Combined Multi-head Latent Attention with fine-grained MoE routing to reduce training and inference cost.
Opened Meta’s new open-model generation in 8B and 70B sizes.
The flagship Claude 3 model, combining visual input with a 200K context window.
Google’s first lightweight open models, released in 2B and 7B sizes.
Used an MoE architecture to extend usable context to 1M tokens while retaining cross-modal retrieval.
Significantly improved long-prompt understanding, coherence, knowledge, and image prompting.
The first Gemini flagship, designed from training onward for native text, image, audio, and video understanding.
The starting point of Qwen’s open-model path, emphasizing Chinese-English, multilingual, and tool-use capabilities.
Used a base-plus-refiner pipeline to bring open generation to native 1024 resolution.
Expanded to a 100K context window and reached broader users through the API and claude.ai.
A large multimodal model for complex professional tasks; core scale and training details were not disclosed.
Anthropic’s first public Claude model, centered on its Constitutional AI training approach.
Brought strong text-to-image weights to local GPUs and catalyzed a broad tooling and fine-tuning ecosystem.
A 175B autoregressive model that systematically demonstrated few-shot and in-context learning without fine-tuning.
Demonstrated broad task transfer from unsupervised language modeling and introduced a staged-weight release.