Overview
An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.
An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.
Only published specifications are shown.
An “omni” model natively spanning text, vision, and audio with substantially lower interaction latency.
GPT-4o moved native multimodality from a research capability into a real-time mass-market experience.
GPT-4o launched with native text, vision, and audio.
View sourceSelect a card to view the related model.