OpenAI resets GPT-5-Codex users’ limits after an expansion slowdown
After a slowdown during additional GPU-capacity deployment, the team resets limits for all GPT-5-Codex users as compensation.
SEARCH
Browse published English reports with original sources and content provenance.
After a slowdown during additional GPU-capacity deployment, the team resets limits for all GPT-5-Codex users as compensation.
An experimental Alibaba release exploring novel routing strategies that sharply lower the active-parameter ratio per forward pass.
Mistral releases Magistral Small 1.2, a maintenance update to its 24-billion-parameter Magistral Small reasoning model.
Mistral releases Magistral Medium 1.2, a closed model tuned for throughput and queue handling in high-concurrency enterprise API workloads.
OpenAI brings the Realtime API to general availability and releases gpt-realtime for production voice agents.
The fourth-generation conversational model from community developers, instruction-tuned on a large open-source base.
Google released Gemini 2.5 Flash Image, code-named Nano Banana, for image generation, editing, and multi-image fusion.
An architecture tune of V3 that improves MoE expert load-balancing and the stability of very large-scale distributed training.
Z.ai releases GLM-4.5V, a 106-billion-parameter visual-language model with stronger image-text alignment and spatial-coordinate understanding.
OpenAI's fifth-generation flagship, introducing a new architecture and setting a new bar across general-purpose tasks.
OpenAI's first hundred-billion-parameter foundation model released to the open-source community.
Anthropic releases Claude Opus 4.1, an incremental closed-model update tuned for complex mathematics and long reasoning chains.
Z.ai releases GLM-4.5 Air, a lighter 106-billion-parameter MoE variant optimized to reduce redundancy and improve inference speed.
Z.ai releases GLM-4.5, a 355-billion-parameter MoE model with updated expert routing and improved long-form structure generation.
ByteDance Seed releases the end-to-end speech-to-speech interpretation model Seed LiveInterpret 2.0 and says it is publicly available through Volcano Engine.
Alibaba Cloud releases Qwen3-Coder, a 480-billion-parameter MoE model for complex software engineering and long-form code generation.
Google releases the stable Gemini 2.5 Flash-Lite model for high-throughput workloads that prioritize low cost and low latency.
ChatGPT agent unifies web operation, deep research, code execution, and editable deliverables into a single agent.
NVIDIA releases a 2.5B English speech-augmented language model combining FastConformer and Qwen.
Moonshot AI's trillion-parameter MoE model, continuing its emphasis on lossless recall across a very long context window.