OpenAI releases GPT-4o, bringing real-time voice conversation to everyone
Event Summary
GPT-4o (May 2024) made real-time voice conversations with AI available to everyone. OpenAI’s omnimodal model combined text, vision, and audio in one neural network.
Impact Assessment
-
Capability Leap +2 · Medium-term
First production model to natively fuse text, vision, and audio in a single neural network. Real-time voice conversation with emotional expression reached human-like quality for the first time. Audio latency of 232ms was indistinguishable from human conversation.
Affected Groups: AI users, developers, accessibility communities
-
Access Democratization +2 · Immediate
GPT-4o-level intelligence became free for all ChatGPT users. API pricing was cut 50% compared to GPT-4 Turbo. The combination of free access + voice interface made state-of-the-art AI conversational ability available to anyone with a smartphone.
Affected Groups: general public, developers, small businesses, students
Consensus & Sources
-
1
We're introducing GPT-4o, our new flagship model that can reason across audio, vision, and text in real time.Reference Evidence Citation logged Live source
-
2
GPT-4o (omni) is a multimodal language model developed by OpenAI and released on May 13, 2024.Reference Evidence Citation logged Live source