Model index
OpenAI·GPT

GPT-4o

An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.

LegacyLanguageMultimodalWeights not released
MODEL NUMBER004
Release date
2024-05-13
Organization
OpenAI
Family
GPT
01

Specifications

Only published specifications are shown.

Architecture
Undisclosed
Context
128K
Input modalities
Text · Image · Audio
Open weights
No
Status
Legacy
02

Overview

Overview

An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.

Why it mattered

GPT-4o moved native multimodality from a research capability into a real-time mass-market experience.

03

Changes from predecessor

Compared with GPT-4

PreviousGPT-4 · 8K ctx
CurrentGPT-4o · 128K ctx
New capabilities Native multimodal Real-time audio
04

Architecture & capabilities

Key capabilities

Native multimodalVisionReal-time audio
05

Release history

General release

GPT-4o launched with native text, vision, and audio.

View source
06

Related models

Select a card to view the related model.

07

Sources