OpenAI releases GPT-4o
OpenAI releases the native multimodal GPT-4o model for real-time interaction with text, images and audio.
ORGANIZATION
English reports linked to this entity will appear here as their translations are published.
OpenAI releases the native multimodal GPT-4o model for real-time interaction with text, images and audio.
OpenAI releases TTS-1 HD for higher-quality speech synthesis at its first DevDay.
At its first DevDay, OpenAI opens DALL·E 3 through the Images API, allowing developers to integrate text-to-image generation under the dall-e-3 model name.
OpenAI announced GPT-4 Turbo with a 128K context window and opened the preview to paying API developers, alongside a separate vision preview model.
OpenAI releases the real-time text-to-speech model TTS-1 at its first DevDay.
OpenAI makes DALL·E 3 available to ChatGPT Plus and Enterprise users, who can generate an image in conversation and request further changes.
OpenAI announces DALL·E 3; the accompanying paper attributes much of its stronger prompt following to recaptioning training images with detailed synthetic descriptions.
OpenAI launches GPT-4 as a model that accepts text and images and emits text; text access begins through ChatGPT Plus and an API waitlist, while image input remains in limited partner testing.
OpenAI makes ChatGPT freely available as a research preview, using a GPT-3.5-series model trained with supervised dialogue data and reinforcement learning from human feedback.
OpenAI makes DALL·E image generation and editing available to all API customers for integration into applications and products.
OpenAI removes the DALL·E beta waitlist, allowing users to sign up immediately; more than 1.5M active users were generating over 2M images per day.
OpenAI released and open-sourced Whisper, an automatic speech recognition model for multilingual transcription, language identification, and speech translation into English.
OpenAI introduces DALL·E 2, built around CLIP image embeddings and a diffusion decoder, and initially opens it only to a small group of invited users.
OpenAI releases a model trained with reinforcement learning from human feedback to better follow user intent.
OpenAI releases a code-generation model fine-tuned from GPT-3 and used in the early GitHub Copilot.
OpenAI introduces DALL·E, a 12B autoregressive Transformer that models text and discrete image tokens as one sequence.
OpenAI released CLIP, a vision-language model trained with natural-language supervision, together with its paper, code, and pretrained models.
OpenAI launches a private beta for a general-purpose text-in, text-out API running GPT-3 family models, offering controlled access instead of open weights.
OpenAI publishes GPT-3 research showing that a 175B-parameter model can switch tasks from instructions and a few examples in the prompt without updating its weights.
OpenAI releases code and model weights for the largest GPT-2 configuration, concluding its nine-month staged-release experiment.