Zhipu AI releases the bilingual ChatGLM-130B
Zhipu AI releases a 130-billion-parameter dialogue model built for Chinese and English, establishing the first generation of the ChatGLM line.
EVENT ARCHIVE
Browse published English reports with original sources and content provenance.
Zhipu AI releases a 130-billion-parameter dialogue model built for Chinese and English, establishing the first generation of the ChatGLM line.
Google Research announces Imagen, a text-to-image diffusion model focused on photorealism and language understanding; its code and public demo were not released at launch.
Meta releases the 125M-to-175B OPT family; smaller checkpoints and code are available, while OPT-175B is limited to approved noncommercial research access.
DeepMind introduced Flamingo, an 80-billion-parameter visual language model that used a few examples to handle image, video, and text tasks.
EleutherAI scales an open foundation model to the 20-billion-parameter range.
OpenAI introduces DALL·E 2, built around CLIP image embeddings and a diffusion decoder, and initially opens it only to a small group of invited users.
Google presents a 540B-parameter dense language model trained with Pathways; the initial release is a research disclosure, without public weights or a general-purpose API.
DeepMind introduced the code-generation system AlphaCode, which performed at about the median level of participants in Codeforces competitions.
The paper documents LaMDA as a dialogue-focused model family with up to 137B parameters and examines response quality, safety and factual grounding.
DeepMind introduced Gopher, a 280-billion-parameter language model, and published research on its capabilities, limitations, and risks without releasing public weights or an API.
OpenAI releases a code-generation model fine-tuned from GPT-3 and used in the early GitHub Copilot.
AI21 Labs launched the Jurassic-1 language model family and opened its API and web interface through the AI21 Studio open beta.
EleutherAI releases an early open 6-billion-parameter model that sees broad academic use.
Google introduces LaMDA for open-domain dialogue at I/O 2021; the model remains an internal research system without public weights or a general-purpose API.
EleutherAI releases an early open alternative to GPT-3.
Switch Transformer routed each token to a single expert, scaling total sparse parameters to about 1.6 trillion without activating the full model for every token.
OpenAI introduces DALL·E, a 12B autoregressive Transformer that models text and discrete image tokens as one sequence.
OpenAI released CLIP, a vision-language model trained with natural-language supervision, together with its paper, code, and pretrained models.
Meta published the BART denoising sequence-to-sequence model and made code and pretrained weights available through fairseq.
Google researchers introduce GShard automatic sharding and use it to train a multilingual translation MoE with more than 600B total parameters on 2,048 TPU v3 accelerators.