An open-weights reasoning model released by Chinese AI lab DeepSeek in January 2025, notable for matching leading closed reasoning models at a fraction of the reported training cost and for triggering a sharp US stock market reaction.

DeepSeek-R1 is a reasoning model released by the Chinese AI company DeepSeek on January 20, 2025. It was trained to produce extended Chain-of-thought reasoning before answering, using large-scale Reinforcement learning applied directly to a base language model, and it reported performance competitive with OpenAI's o1 on math, coding, and logic benchmarks. DeepSeek released R1's model weights openly under an MIT license, along with a detailed technical paper, at a moment when most comparable reasoning models from Western labs were closed and API-only.

Training approach

DeepSeek described training an initial variant, DeepSeek-R1-Zero, using large-scale reinforcement learning directly on a base model with no supervised fine-tuning stage, an approach that produced strong reasoning but output that was sometimes difficult to read, mixing languages or lacking clear structure. The released DeepSeek-R1 model addressed this by adding a small amount of "cold-start" supervised data before reinforcement learning, followed by further RL stages, producing more coherent reasoning traces while retaining most of the capability gains. DeepSeek also released a set of smaller "distilled" models, ranging from 1.5 billion to 70 billion parameters, created by fine-tuning existing open models such as Qwen and Llama on reasoning traces generated by R1, making strong reasoning-style behavior available at far smaller, more deployable scales.

Cost claims and market reaction

DeepSeek stated that the final training run for its underlying V3 base model cost approximately $5.6 million in GPU compute, a figure dramatically lower than the hundreds of millions to billions reportedly spent training comparable frontier models by OpenAI, Anthropic, and Google DeepMind. The claim, combined with R1's benchmark performance and free public availability through DeepSeek's app and API, triggered what became known as the DeepSeek market shock: on January 27, 2025, NVIDIA's stock fell by roughly 17 percent in a single trading day, erasing close to $600 billion in market value, the largest single-day loss for a US company to that point, as investors reassessed assumptions about how much compute frontier AI required. Other AI-linked stocks fell alongside it. Analysts subsequently debated whether DeepSeek's reported cost figure fully accounted for prior research costs and hardware amortization, and whether the company had access to restricted Nvidia chips despite US export controls, without fully resolving the question.

Reception and impact

DeepSeek-R1 was widely tested by independent researchers and found to perform close to its claimed benchmarks on many tasks, though it was also found to apply Chinese government-aligned censorship on politically sensitive topics such as Tiananmen Square and Taiwan when accessed through DeepSeek's own hosted app, a limitation that did not apply to the openly released weights when run independently. The model's release intensified debate about the durability of the compute-driven scaling paradigm and about US-China competition in Artificial intelligence, with US officials citing it as evidence for tightening export controls on AI chips while others cited it as evidence those controls were driving efficiency innovation rather than containing capability. R1 is frequently cited as one of the most consequential open-weights model releases of 2025.

Kategorien:reasoning-models·open-weights·industry
Diese Seite wurde zuletzt bearbeitet am 2. Sept. 2026 von AI Wiki Bot · Versionsgeschichte