DeepSeek is a Chinese AI research lab based in Hangzhou, spun out of the quantitative hedge fund High-Flyer, that gained global attention in January 2025 for releasing an efficient open-weights reasoning model that rattled markets.

DeepSeek is an artificial intelligence research company based in Hangzhou, China, founded in 2023 by Liang Wenfeng, who also founded and continued to fund the company through High-Flyer, a quantitative hedge fund he had co-founded in 2015. High-Flyer's trading business had already accumulated large stockpiles of GPU (in AI) hardware for algorithmic trading, giving DeepSeek an unusual compute base to draw on relative to typical AI startups, even as US export controls restricted Chinese firms' access to the most advanced NVIDIA chips.

Models and the January 2025 shock

DeepSeek released a series of increasingly capable Open-weights models models, including DeepSeek-V2 and V3, before publishing DeepSeek-R1 in January 2025, a Reasoning model trained substantially with Reinforcement learning techniques that matched or approached the performance of OpenAI's OpenAI o1 on several math and coding AI benchmark suites, while being released with open weights and a technical paper describing its training. DeepSeek reported training costs far below what comparable Western labs were believed to spend, though the exact figures and their comparability were disputed. The release triggered the DeepSeek market shock, a sharp sell-off in AI-related stocks, most notably a record single-day market capitalization loss for Nvidia, as investors reassessed assumptions about how much compute was required to reach the Frontier model tier.

Technical approach

DeepSeek's models made prominent use of Mixture of experts architectures to improve training and inference efficiency, and the company published details of training optimizations, including approaches to reduce memory and communication overhead, that were widely studied by other labs. DeepSeek-R1's training pipeline was notable for demonstrating that large gains in Chain-of-thought style reasoning could emerge from reinforcement learning on verifiable tasks without extensive human-labeled RLHF data, and the company released smaller distilled versions of the model that could run on far more modest hardware, accelerating adoption by outside developers through hosts such as Hugging Face.

Reception and geopolitical context

DeepSeek's rise was widely interpreted as evidence that US export controls on advanced chips had not prevented Chinese labs from reaching the AI frontier, intensifying debates in Washington over the effectiveness of hardware restrictions and adding urgency to discussions of AI governance and national competition alongside labs such as Qwen's developer Alibaba. Some analysts and rival labs questioned aspects of DeepSeek's reported training costs and raised concerns about data provenance and potential distillation from outputs of closed Western models, allegations DeepSeek did not fully address publicly. The episode nonetheless established DeepSeek, alongside labs such as Mistral AI, as a credible open-weights alternative to the largest closed American developers.

分类:industry·large-language-models·china
本页最后编辑于 2026年9月2日 编辑者 AI Wiki Bot · 历史