DeepSeek-R1 is an open-weights large language model developed by DeepSeek, a Chinese artificial intelligence company based in Hangzhou, Zhejiang. Released in January 2025 alongside the company's eponymous chatbot, DeepSeek-R1 gained international attention for its advanced reasoning abilities and for shaking global financial markets, as its low-cost development challenged assumptions about the dominance of Western AI firms. The model is part of DeepSeek's broader strategy of publishing open-weight models under free and open-source software licenses, making its parameters publicly available while keeping training data proprietary.
DeepSeek was founded in July 2023 by Liang Wenfeng, who also co-founded the hedge fund High-Flyer, which owns and funds the company. The release of DeepSeek-R1 in January 2025 marked a significant milestone in the company's history, positioning it as a major player in the global AI landscape. The model's performance, combined with its efficient training methods, sparked debates about the future of AI development, particularly regarding the role of open-source models and the impact of US chip export restrictions on Chinese AI innovation.
Background
DeepSeek traces its origins to High-Flyer, a Chinese hedge fund co-founded by Liang Wenfeng in June 2015. Liang, who had been trading since the 2008 financial crisis while attending Zhejiang University, began using GPU-dependent deep learning models for stock trading on 21 October 2016, replacing earlier CPU-based linear models. By the end of 2017, most of High-Flyer's trading was driven by AI.
In 2019, High-Flyer constructed its first computing cluster, Fire-Flyer, at a cost of 200 million yuan. The cluster contained 1,100 GPUs interconnected at 200 Gbit/s and was retired after 1.5 years of operation. By 2021, Liang had begun purchasing large quantities of Nvidia GPUs for an AI project, reportedly obtaining 10,000 Nvidia A100 GPUs before the United States restricted chip sales to China.
Construction of Fire-Flyer 2 began in 2021 with a budget of 1 billion yuan. In 2022, Fire-Flyer 2's capacity was used at over 96%, totaling 56.74 million GPU hours, with 27% of its capacity supporting scientific computing outside the company. The cluster had 5,000 PCIe A100 GPUs in 625 nodes, each containing 8 GPUs, and later incorporated NVLinks and Nvidia Collective Communications Library (NCCL) to train larger models requiring model parallelism.
Founding and Early Development
On 14 April 2023, High-Flyer announced the launch of an artificial general intelligence (AGI) research lab, stating that the new lab would focus on developing AI tools unrelated to the firm's financial business. Two months later, on 17 July 2023, that lab was spun off into an independent company named DeepSeek, with High-Flyer as its principal investor and backer. Initially, venture capital investors were reluctant to provide funding, as they considered it unlikely that the venture would be able to quickly generate an exit.
DeepSeek is headquartered in Hangzhou, Zhejiang, and is owned and funded by High-Flyer. Its co-founder, Liang Wenfeng, serves as CEO. As of May 2024, Liang personally held an 84% stake in DeepSeek through two shell corporations. The company's strategy emphasizes research over immediate commercialization, allowing it to skirt certain provisions of China's AI regulations aimed at consumer-facing technologies.
Release of DeepSeek-R1
DeepSeek-R1 was released in January 2025, alongside the launch of the company's eponymous chatbot. The model's release generated significant media coverage and market reactions, as its performance rivaled that of leading Western models such as those from OpenAI and Anthropic, while reportedly being trained at a fraction of the cost. The release was seen as a demonstration of China's growing capabilities in artificial intelligence and a challenge to the assumption that advanced AI development required massive computational resources.
The model's open-weights nature meant that its parameters were publicly shared, allowing researchers and developers worldwide to download and fine-tune it. This contrasted with the proprietary approaches of many Western AI companies, which kept their models' weights confidential. DeepSeek-R1's release also highlighted the company's ability to innovate under US chip export restrictions, as it had refined its algorithms to maximize computational efficiency using older hardware.
Technical and Operational Aspects
DeepSeek-R1 is built on the transformer architecture, the foundation of most modern large language models. The model's training leveraged the Fire-Flyer 2 computing cluster, which consists of co-designed software and hardware architecture. The hardware side uses Nvidia GPUs with 200 Gbps interconnects, divided into two zones supporting cross-zone tasks. The network topology consists of two fat trees, chosen for their high bisection bandwidth.
The software side includes 3FS (Fire-Flyer File System), a distributed parallel file system designed for asynchronous random reads using Direct I/O and RDMA Read. Unlike standard Buffered I/O, Direct I/O does not cache data, which is beneficial since each piece of data read is random and not reused. The cluster also uses hfreduce, an asynchronous communication library originally designed to replace Nvidia Collective Communications Library (NCCL) for certain operations.
DeepSeek's training approach emphasizes efficiency, allowing the company to achieve competitive performance despite hardware constraints. The company has stated that it focuses on research and does not have immediate plans for commercialization, which has enabled it to prioritize technical innovation over market pressures.
Market Impact and Reception
The release of DeepSeek-R1 in January 2025 had a significant impact on global financial markets, particularly on technology stocks. The model's performance, combined with reports of its low training costs, led to a sell-off in shares of companies perceived as leaders in AI infrastructure, such as Nvidia and other chipmakers. The event was widely interpreted as a sign that the competitive landscape of AI was shifting, with open-source models from China potentially disrupting the dominance of US-based AI companies.
DeepSeek-R1 received praise for its reasoning capabilities, which were comparable to or exceeded those of leading proprietary models in certain benchmarks. The model's open-weights nature also facilitated its adoption by cloud providers, including Microsoft Azure and Perplexity AI, which hosted it natively. This broad availability contributed to its popularity among researchers and developers.
However, the release also raised concerns about the implications of open-weight models for AI safety and regulation. Some experts argued that the widespread availability of powerful reasoning models could pose risks, while others welcomed the democratization of AI technology. The event sparked discussions about the need for international cooperation on AI governance, particularly in light of the geopolitical tensions between the US and China.
Broader Context and Future Directions
DeepSeek-R1's release is part of a broader trend of Chinese AI companies developing competitive models despite US export controls. DeepSeek has continuously refined its algorithms to maximize computational efficiency, leveraging older hardware and reducing energy consumption. The company has also expanded into Africa, offering more affordable and less power-hungry AI solutions, and has bolstered African language models, generating a number of startups, for example in Nairobi. Along with Huawei's storage and cloud computing services, the impact on the tech scene in sub-Saharan Africa is considerable, with DeepSeek offering local data sovereignty and more flexibility compared to Western AI platforms.
The company's hiring approach emphasizes skills over lengthy work experience, resulting in many hires fresh out of university. DeepSeek also recruits individuals without computer science backgrounds to expand the range of expertise incorporated into the models, for instance in poetry or advanced mathematics. According to The New York Times, dozens of DeepSeek researchers have or have previously had affiliations with People's Liberation Army laboratories and the Seven Sons of National Defence. Since 2025, its chatbot was adopted in non-combat roles by the People's Liberation Army, including hospitals, the People's Armed Police paramilitary, and national mobilization organizations.
As of early 2025, DeepSeek-R1 remains a significant reference point in the AI community, illustrating the potential of open-source development and the increasing global competition in generative AI. The model's release has prompted other companies to reconsider their strategies, with some exploring more open approaches to model distribution. The long-term impact of DeepSeek-R1 on the AI industry is still unfolding, but its immediate effect on markets and public discourse has been profound.