DeepMind introduces the Flamingo visual language model
DeepMind introduced Flamingo, an 80-billion-parameter visual language model that used a few examples to handle image, video, and text tasks.
What happened
Event details
DeepMind introduced Flamingo on April 28, 2022. The model accepted prompts containing interleaved images, videos, and text and generated text responses. It connected pretrained visual and language models to perform multimodal few-shot learning. The final model described by DeepMind had 80 billion parameters, but no weights or public access endpoint were released.
Assessment
Why it matters
Flamingo showed that one visual language model could use a small number of examples to handle many image and video tasks, making it an important step toward general-purpose multimodal models.
Availability
Access notes
The research was public, but the model itself was not available for public use.