Mistral AI releases Mamba Codestral 7B
Mistral AI releases a code-generation model built with a non-Transformer Mamba architecture.
EVENT ARCHIVE
20 English reports are shown on this page. Refine the keywords, date range, company, or event type below.
Mistral AI releases a code-generation model built with a non-Transformer Mamba architecture.
NVIDIA releases a 340-billion-parameter open model and supporting materials for synthetic-data generation.
Stability AI released the 2-billion-parameter Stable Diffusion 3 Medium with downloadable model weights and API access.
Alibaba Cloud releases the second Qwen generation with stronger multilingual support and long-text handling.
Z.ai releases the fourth GLM generation with substantially stronger overall benchmark results.
Mistral AI releases an open model optimized for code generation and programming tasks.
At Google I/O 2024, Google announces Imagen 3 with improvements in photorealism, complex prompt understanding, detail and text rendered in images.
OpenAI releases the native multimodal GPT-4o model for real-time interaction with text, images and audio.
DeepSeek releases a 236B mixture-of-experts model that combines sparse activation with latent attention to cut training and inference costs.
Meta FAIR researchers propose training language models to predict several future tokens at once, work that later informs DeepSeek-V3's MTP design.
Snowflake releases a mixture-of-experts model optimized for SQL queries and enterprise code generation.
Microsoft releases the third Phi generation, optimized for efficient execution on edge devices.
Meta releases its fourth foundation-model generation with more training data and stronger same-scale performance.
Mistral AI releases an open mixture-of-experts model with a 176-billion-parameter scale.
Reka AI releases a multimodal model that accepts text, image, video and audio inputs together.
Microsoft releases a new open model family with strong instruction following and complex-task performance.
Cohere releases an enterprise model optimized for retrieval-augmented generation and tool use.
AI21 Labs releases a hybrid model combining Mamba and Transformer architectures.
Databricks releases an open 132-billion-parameter mixture-of-experts model that raises open-model benchmark performance.
xAI released the 314-billion-parameter Grok-1 base-model weights and network architecture under Apache 2.0.