ByteDance publishes Seed-TTS
ByteDance Seed publishes the Seed-TTS speech-generation foundation model with voice cloning and controllable synthesis capabilities.
EVENT ARCHIVE
English reports are listed here. AI-generated drafts are clearly labeled until an editor reviews them.
ByteDance Seed publishes the Seed-TTS speech-generation foundation model with voice cloning and controllable synthesis capabilities.
Claude 3.5 Sonnet improves reasoning, coding and vision, and introduces Artifacts.
Runway releases Gen-3 Alpha with improvements in video-generation fidelity, temporal consistency, motion and camera control.
DeepSeek releases a V2 architecture model optimized for programming and mathematical reasoning.
Mistral AI releases a code-generation model built with a non-Transformer Mamba architecture.
NVIDIA releases a 340-billion-parameter open model and supporting materials for synthetic-data generation.
Stability AI released the 2-billion-parameter Stable Diffusion 3 Medium with downloadable model weights and API access.
Alibaba Cloud releases the second Qwen generation with stronger multilingual support and long-text handling.
Z.ai releases the fourth GLM generation with substantially stronger overall benchmark results.
Mistral AI releases an open model optimized for code generation and programming tasks.
At Google I/O 2024, Google announces Imagen 3 with improvements in photorealism, complex prompt understanding, detail and text rendered in images.
OpenAI releases the native multimodal GPT-4o model for real-time interaction with text, images and audio.
DeepSeek releases a 236B mixture-of-experts model that combines sparse activation with latent attention to cut training and inference costs.
Meta FAIR researchers propose training language models to predict several future tokens at once, work that later informs DeepSeek-V3's MTP design.
Snowflake releases a mixture-of-experts model optimized for SQL queries and enterprise code generation.
Microsoft releases the third Phi generation, optimized for efficient execution on edge devices.
Meta releases its fourth foundation-model generation with more training data and stronger same-scale performance.
Mistral AI releases an open mixture-of-experts model with a 176-billion-parameter scale.
Reka AI releases a multimodal model that accepts text, image, video and audio inputs together.
Microsoft releases a new open model family with strong instruction following and complex-task performance.