DeepSeek's low-cost training playbook has become the industry's default assumption, not an anomaly
The methods DeepSeek popularised for training highly capable models at a fraction of the compute cost that Western frontier labs were reporting - aggressive mixture-of-experts architectures, reinforcement-learning-driven reasoning training, and engineering optimisations that squeeze more effective compute out of a given chip allocation - have moved from a startling one-off disclosure into the default set of techniques that essentially every serious model-training team, Chinese and Western alike, now assumes as baseline practice. The immediate market shock the original disclosure caused, including a sharp but temporary selloff in AI-infrastructure stocks on fears that massive compute spending might be unnecessary, has given way to a more settled industry consensus: efficiency gains lower the cost of achieving a given capability level, but the labs with the most compute still push the capability frontier furthest, meaning the compute arms race and the efficiency race are complementary rather than substitutes for each other. Chinese open-weight models broadly - spanning DeepSeek's continued releases, Alibaba's Qwen series, and others - have continued to close the capability gap with the best closed American models on many benchmarks while remaining freely downloadable, a dynamic that has forced OpenAI, Anthropic and Google to justify their pricing and API-access models increasingly on enterprise trust, safety guarantees, tooling ecosystem and support rather than on raw model capability alone, since a meaningful share of that capability is now available to any developer willing to self-host an open-weight alternative. US export-control policy toward Chinese AI development has continued to grapple with the reality that DeepSeek-style efficiency gains blunt the intended effect of chip export restrictions - if leading Chinese labs can achieve competitive results with fewer or lower-end chips than restrictions assume are necessary, the policy's central premise weakens, and successive rounds of export-control refinement have tried, with uneven success, to close the gaps that efficiency innovations keep reopening. For cost-sensitive AI application builders globally, including a large cohort of Indian startups building on top of foundation models rather than training their own, the proliferation of highly capable open-weight models has been an unambiguous benefit, lowering the cost of building AI-native products and reducing dependence on any single closed-model provider's pricing and API terms. What to watch: whether the next generation of Chinese open-weight releases closes the remaining gap on the hardest reasoning and agentic benchmarks, how US export-control policy adapts to efficiency-driven circumvention of chip-based restrictions, and whether any Western lab responds by open-weighting a frontier-class model of its own to compete on the same terms.
Original source: Wired