Xiaomi releases MiMo V2.5 multimodal agent model
Xiaomi releases MiMo V2.5 for native text, image, video and audio understanding; model weights follow under the MIT license on June 29.
ORGANIZATION
English reports linked to this entity will appear here as their translations are published.
Xiaomi releases MiMo V2.5 for native text, image, video and audio understanding; model weights follow under the MIT license on June 29.
Xiaomi releases MiMo V2.5 Pro for complex agent and coding tasks, with a 1M-token context window and MIT-licensed weights from June 29.
Xiaomi launches its flagship MiMo V2 Pro through the API, with a 1M-token context window and hybrid attention for complex agent workflows.
Xiaomi launches MiMo V2 Omni through the API, unifying text, vision and speech perception with tool execution and a 256K context window.
Xiaomi open-sourced an efficient MoE base for code and agents under an MIT license, using hybrid attention and a multi-layer MTP design.
Xiaomi releases the SFT and RL versions of the open MiMo VL 7B vision-language model for general visual understanding and multimodal reasoning.
Xiaomi releases the open MiMo 7B family for mathematics, coding and general reasoning, covering Base, SFT, RL-Zero and RL variants.