Model releaseHigh global significanceConfirmed confidence
OpenAI releases GPT-4o Transcribe Diarize
OpenAI releases an ASR model with built-in speaker diarization and support for reference speaker clips.
What happened
Event details
The model combines transcription, speaker diarization and segment timestamps in the Transcription API, and can use reference audio to improve labels for known speakers.
Assessment
Why it matters
It makes diarization a native capability of a specialized ASR model and adds reference-audio-assisted labeling for known speakers.
72/100Global significance score. Regional effects are recorded only when the evidence supports a meaningful difference.
Availability
Access notes
Available only through the Transcription API. Users can provide 2–10 seconds of reference audio for up to four known speakers. The release date is the model's first official documented availability.