Back to events
Model releaseHigh global significanceConfirmed confidence

OpenAI releases GPT-4o Transcribe Diarize

OpenAI releases an ASR model with built-in speaker diarization and support for reference speaker clips.

Event details

The model combines transcription, speaker diarization and segment timestamps in the Transcription API, and can use reference audio to improve labels for known speakers.

Why it matters

It makes diarization a native capability of a specialized ASR model and adds reference-audio-assisted labeling for known speakers.

72/100Global significance score. Regional effects are recorded only when the evidence supports a meaningful difference.

Access notes

Available only through the Transcription API. Users can provide 2–10 seconds of reference audio for up to four known speakers. The release date is the model's first official documented availability.