OpenAI merges Operator and deep research into ChatGPT agent
ChatGPT agent unifies web operation, deep research, code execution, and editable deliverables into a single agent.
EVENT ARCHIVE
English reports are listed here. AI-generated drafts are clearly labeled until an editor reviews them.
ChatGPT agent unifies web operation, deep research, code execution, and editable deliverables into a single agent.
Alongside the Claude 4 launch, Anthropic moved Claude Code out of research preview to general availability across terminal, IDE, background tasks, and SDK.
Google released Veo 3, bringing native generation of ambient sound, sound effects, and character dialogue to its flagship video model.
OpenAI released Codex, a software-engineering agent that handles coding tasks in parallel inside a cloud sandbox.
OpenAI opened ChatGPT's image generation to developers as the gpt-image-1 model in the Images API.
Manus asks users for an outcome, then plans, researches, codes, operates websites, and returns finished artifacts from a cloud virtual machine.
Hugging Face launches an open community project to reproduce R1 and fill in DeepSeek's unpublished training data and complete training pipeline.
DeepSeek passes ChatGPT on the U.S. App Store's free chart as Nvidia loses roughly $593 billion in market value during an AI-sector selloff.
Operator uses a visual computer-use model to click, type, and scroll through web tasks, handing control back for logins, payments, and sensitive actions.
DeepSeek publishes R1, R1-Zero, six distilled models, and its reinforcement-learning recipe, pushing open reasoning models into the frontier conversation.
DeepSeek puts V3, web search, a reasoning mode, and file analysis into a free mobile app five days before the R1 release.
DeepSeek releases a 671B MoE model with 37B active parameters, open weights, low API prices, and unusually detailed efficiency disclosures.
OpenAI moves Sora from research preview to a standalone video-generation product for ChatGPT Plus and Pro users.
Tencent released HunyuanVideo, a text-to-video model with more than 13 billion parameters, together with model weights and inference code.
Meta released Llama 3.2 with 1B and 3B lightweight text models and 11B and 90B vision-language models.
Black Forest Labs releases the FLUX.1 text-to-image model family with an API version and open-weight versions under different licenses.
Mistral AI releases a code-generation model built with a non-Transformer Mamba architecture.
OpenAI releases the native multimodal GPT-4o model for real-time interaction with text, images and audio.
Meta FAIR researchers propose training language models to predict several future tokens at once, work that later informs DeepSeek-V3's MTP design.
AI21 Labs releases a hybrid model combining Mamba and Transformer architectures.