CURATED TIMELINE · EDITORIAL EDITION

DeepSeek Before the Breakout

The January 2025 sensation did not appear overnight. Trace the open-weight code models, sparse experts, GRPO, MTP, V2 efficiency work, R1 preview, V3, the free app, and the App Store shock — while separating reported training-run cost from total R&D and open weights from full open source.

Timeline overview

At a glance

Hover for the event briefScroll horizontally
10 nodes2023.11—2025.01 · Editorially curated · Updated 2026-07-15

Editorial thread

The timeline

Reviewed event briefs and original editorial context, ordered to show how the story changed over time.

  1. 01
    Paper & weights release

    DeepSeek publishes the sparse architecture behind its later models

    DeepSeekMoE uses fine-grained experts and shared expert isolation to improve sparse-model efficiency, with 16B weights and training code released.

  2. 02
    Open-source release

    DeepSeek starts with an open-weight code model

    DeepSeek releases code models from 1.3B to 33B parameters, making open weights and code part of its product strategy from the outset.