Back to Timeline
2024-05-13

OpenAI releases GPT-4o, bringing real-time voice conversation to everyone

Capability Breakthrough

Event Summary

GPT-4o (May 2024) made real-time voice conversations with AI available to everyone. OpenAI’s omnimodal model combined text, vision, and audio in one neural network.

Impact Assessment

  • Capability Leap +2 · Medium-term

    First production model to natively fuse text, vision, and audio in a single neural network. Real-time voice conversation with emotional expression reached human-like quality for the first time. Audio latency of 232ms was indistinguishable from human conversation.

    Affected Groups: AI users, developers, accessibility communities

  • Access Democratization +2 · Immediate

    GPT-4o-level intelligence became free for all ChatGPT users. API pricing was cut 50% compared to GPT-4 Turbo. The combination of free access + voice interface made state-of-the-art AI conversational ability available to anyone with a smartphone.

    Affected Groups: general public, developers, small businesses, students

Consensus & Sources

Significance L1
Category Capability Breakthrough
Consensus Broad Consensus
Impact Index 6/10
  • 1

    URL: https://openai.com/index/hello-gpt-4o/

    We're introducing GPT-4o, our new flagship model that can reason across audio, vision, and text in real time.
    Reference Evidence Citation logged Live source
  • 2

    URL: https://en.wikipedia.org/wiki/GPT-4o

    GPT-4o (omni) is a multimodal language model developed by OpenAI and released on May 13, 2024.
    Reference Evidence Citation logged Live source