DeepSeek publishes the sparse architecture behind its later models
DeepSeekMoE uses fine-grained experts and shared expert isolation to improve sparse-model efficiency, with 16B weights and training code released.
CURATED TIMELINE · EDITORIAL EDITION
The January 2025 sensation did not appear overnight. Trace the open-weight code models, sparse experts, GRPO, MTP, V2 efficiency work, R1 preview, V3, the free app, and the App Store shock — while separating reported training-run cost from total R&D and open weights from full open source.
Timeline overview
Editorial thread
Reviewed event briefs and original editorial context, ordered to show how the story changed over time.
DeepSeekMoE uses fine-grained experts and shared expert isolation to improve sparse-model efficiency, with 16B weights and training code released.
DeepSeek releases code models from 1.3B to 33B parameters, making open weights and code part of its product strategy from the outset.
Efficiency became a research program. DeepSeekMoE split experts more finely and isolated shared knowledge so only a small slice of a large model had to run for each token. That line later fed directly into V2 and V3.
Export controls are part of the operating context, but public evidence does not prove that each architectural idea was caused by them. The defensible claim is that efficiency became central under constrained hardware conditions.
The origin story is developer credibility, not an overnight consumer hit. High-Flyer supplied capital and early compute, but DeepSeek operated as a dedicated AI research effort. Calling it a quant fund's casual side project understates the sustained work that followed.
Its first model set the distribution strategy: release weights, code, and technical material to developers before building a mass-market chat product.