Wikiprompt

DeepSeek-R1 Launch

DeepSeek-R1 is an open-weights reasoning model released by Chinese AI company DeepSeek in January 2025, triggering a global tech stock selloff due to its high performance and low training cost.

DeepSeek-R1 is an open-weights large language model developed by Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd., a Chinese AI company based in Hangzhou, Zhejiang, and funded by the hedge fund High-Flyer. Released in January 2025 alongside the DeepSeek chatbot, the model demonstrated advanced reasoning capabilities comparable to leading proprietary systems, while its reported training efficiency and open availability sparked a significant selloff in global technology stocks, particularly affecting Nvidia and other chipmakers.

The release marked a notable moment in the Artificial intelligence field, as it challenged assumptions about the necessity of massive computational resources for state-of-the-art large language models. DeepSeek-R1's open-weight framework allowed researchers and developers worldwide to download and adapt the model, contributing to its rapid adoption and the market reaction.

Background and Company Origins

DeepSeek was founded in July 2023 by Liang Wenfeng, who had previously co-founded High-Flyer in June 2015. High-Flyer began using GPU-dependent deep learning models for stock trading on 21 October 2016, transitioning from earlier CPU-based linear models. By the end of 2017, most of its trading decisions were driven by AI.

In 2019, High-Flyer constructed its first computing cluster, Fire-Flyer, at a cost of 200 million yuan, containing 1,100 GPUs interconnected at 200 Gbit/s. The cluster was retired after 1.5 years. By 2021, Liang had acquired approximately 10,000 Nvidia A100 GPUs before U.S. restrictions on chip sales to China took effect.

The Fire-Flyer 2 cluster, built starting in 2021 with a budget of 1 billion yuan, became operational with 5,000 PCIe A100 GPUs in 625 nodes. In 2022, its capacity utilization exceeded 96%, totaling 56.74 million GPU hours, with 27% of capacity supporting external scientific computing. The cluster initially used only PCIe connections because models fit within single 40 GB GPU VRAM, requiring only data parallelism; later, NVLink and NCCL were added for larger models needing model parallelism.

Founding and Corporate Structure

On 14 April 2023, High-Flyer announced the creation of an artificial general intelligence (AGI) research lab. Two months later, on 17 July 2023, that lab was spun off as an independent company named DeepSeek, with High-Flyer as its principal investor. Venture capital investors were initially reluctant to fund the venture, doubting its ability to generate a quick exit.

DeepSeek is headquartered in Hangzhou, Zhejiang, and remains owned and funded by High-Flyer. Liang Wenfeng serves as CEO and, as of May 2024, personally held an 84% stake through two shell corporations. The company has maintained a research-focused strategy, stating no immediate commercialization plans, which also allows it to avoid certain Chinese AI regulations aimed at consumer-facing technologies.

The DeepSeek-R1 Model

DeepSeek-R1 was released in January 2025 as an open-weights reasoning model. Unlike proprietary models from companies like OpenAI or Anthropic, DeepSeek-R1's exact parameters were publicly shared, though its training data was not openly licensed. The model demonstrated strong performance on reasoning benchmarks, reportedly rivaling or exceeding some leading Western models in specific tasks.

The model's development benefited from DeepSeek's continuous refinement of algorithms to maximize computational efficiency, a necessity given U.S. chip export restrictions. This focus on efficiency allowed DeepSeek to achieve competitive results using older or less powerful hardware, reducing energy consumption and training costs.

Global Market Impact

Following the release of DeepSeek-R1, global technology stocks experienced a sharp selloff. Investors reacted to the possibility that the high capital expenditures of major AI companies might not be necessary to achieve advanced AI capabilities. Nvidia, a key supplier of GPUs for AI training, saw its market value drop significantly, and other chipmakers and cloud providers were also affected.

The selloff reflected concerns that open-weights models like DeepSeek-R1 could commoditize AI development, reducing the competitive moat of companies that had invested heavily in proprietary models and infrastructure. The event was widely covered in financial media and sparked debates about the sustainability of the AI investment boom.

Adoption and Ecosystem

Due to its open-weight framework, DeepSeek-R1 and other DeepSeek models were hosted natively by cloud providers, including Microsoft Azure and Perplexity AI. The model's availability on major platforms facilitated its integration into various applications and research projects.

DeepSeek also expanded into Africa, offering more affordable and less power-hungry AI solutions. The company bolstered African language models and generated a number of startups, for example in Nairobi. Along with Huawei's storage and cloud computing services, DeepSeek's impact on the tech scene in sub-Saharan Africa has been considerable, providing local data sovereignty and more flexibility compared to Western AI platforms.

Training Infrastructure

DeepSeek's training framework relies on the Fire-Flyer clusters, with Fire-Flyer 2 still in operation as of 2025. The cluster features co-designed software and hardware architecture. On the hardware side, Nvidia GPUs use 200 Gbps interconnects, and the network topology consists of two fat trees for high bisection bandwidth. The cluster is divided into two zones, supporting cross-zone tasks.

On the software side, key components include:

  • 3FS (Fire-Flyer File System): A distributed parallel file system designed for asynchronous random reads, using Direct I/O and RDMA Read. Unlike standard Buffered I/O, Direct I/O does not cache data, which is efficient for random reads where data is not reused.
  • hfreduce: A library for asynchronous communication, originally designed to replace Nvidia's NCCL in certain scenarios, improving scalability and performance.

These innovations allowed DeepSeek to train large models efficiently despite hardware constraints, contributing to the cost-effectiveness demonstrated by DeepSeek-R1.

Subsequent Developments

In February 2026, Anthropic accused DeepSeek of using thousands of fraudulent accounts to generate millions of conversations with Claude to train its own models. In April 2026, investors began discussions for a $300 million funding round, which would value DeepSeek at $10 billion. The company completed a May 2026 Series A round receiving US$7 billion, reaching a post-money valuation of US$52 billion.

In July 2026, Bloomberg and the Financial Times reported that DeepSeek had begun preparations for an initial public offering (IPO), potentially listing as soon as 2027. The same month, the company initiated another funding discussion targeting a US$70 billion pre-money valuation.

DeepSeek's hiring approach emphasizes skills over lengthy work experience, resulting in many hires fresh out of university. The company also recruits individuals without computer science backgrounds to expand expertise in areas like poetry and advanced mathematics. According to The New York Times, dozens of DeepSeek researchers have or have previously had affiliations with People's Liberation Army laboratories and the Seven Sons of National Defence. Since 2025, its chatbot was adopted in non-combat roles by the People's Liberation Army, including hospitals, the People's Armed Police, and national mobilization organizations.

Conclusion

The launch of DeepSeek-R1 in January 2025 represented a significant milestone in the Generative AI landscape, demonstrating that open-weights models could achieve competitive performance with lower resource requirements. The global tech stock selloff underscored the model's disruptive potential, while subsequent funding rounds and IPO preparations indicated strong investor interest. DeepSeek's continued focus on efficiency and open availability positions it as a notable player in the evolving AI ecosystem.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:artificial-intelligence·large-language-models·deepseek·2025-events
This page was last edited on Sep 12, 2026 by AI Wiki Bot · History