OpenAI opens the DALL·E 2 API in public beta
OpenAI makes DALL·E image generation and editing available to all API customers for integration into applications and products.
EVENT ARCHIVE
English reports are listed here. AI-generated drafts are clearly labeled until an editor reviews them.
OpenAI makes DALL·E image generation and editing available to all API customers for integration into applications and products.
Google instruction-tuned the T5 family on 1,836 tasks and released five checkpoints ranging from 80 million to 11 billion parameters.
OpenAI removes the DALL·E beta waitlist, allowing users to sign up immediately; more than 1.5M active users were generating over 2M images per day.
OpenAI released and open-sourced Whisper, an automatic speech recognition model for multilingual transcription, language identification, and speech translation into English.
Stability AI publicly releases Stable Diffusion, bringing an open-weight image-generation model to developers and creator communities.
Midjourney opens public testing and provides a low-barrier text-to-image service through Discord.
DeepMind introduced Flamingo, an 80-billion-parameter visual language model that used a few examples to handle image, video, and text tasks.
OpenAI introduces DALL·E 2, built around CLIP image embeddings and a diffusion decoder, and initially opens it only to a small group of invited users.
DeepMind introduced the code-generation system AlphaCode, which performed at about the median level of participants in Codeforces competitions.
DeepMind introduced Gopher, a 280-billion-parameter language model, and published research on its capabilities, limitations, and risks without releasing public weights or an API.
OpenAI releases a code-generation model fine-tuned from GPT-3 and used in the early GitHub Copilot.
EleutherAI releases an early open alternative to GPT-3.
Switch Transformer routed each token to a single expert, scaling total sparse parameters to about 1.6 trillion without activating the full model for every token.
OpenAI introduces DALL·E, a 12B autoregressive Transformer that models text and discrete image tokens as one sequence.
OpenAI released CLIP, a vision-language model trained with natural-language supervision, together with its paper, code, and pretrained models.
Meta published the BART denoising sequence-to-sequence model and made code and pretrained weights available through fairseq.
Google researchers introduce GShard automatic sharding and use it to train a multilingual translation MoE with more than 600B total parameters on 2,048 TPU v3 accelerators.
OpenAI launches a private beta for a general-purpose text-in, text-out API running GPT-3 family models, offering controlled access instead of open weights.
OpenAI publishes GPT-3 research showing that a 175B-parameter model can switch tasks from instructions and a few examples in the prompt without updating its weights.
T5 cast a wide range of natural-language-processing tasks into a text-to-text format and released a model family ranging from 60 million to 11 billion parameters.