DeepSeek Secretly Developing Custom Inference Chip: AI Arms Race Extends to Silicon
On July 7, 2026, Reuters reported, citing three sources familiar with the matter, that DeepSeek is secretly developing its own AI chip. The chip is designed specifically for inference rather than training, and the project has been underway for approximately one year. Sources indicate that DeepSeek has already engaged with chip design firms, foundries, and memory suppliers, though the project remains in its early stages with no guarantee of success.
This move signals that the Chinese AI company — renowned for its extreme algorithmic efficiency — is officially entering the hardware arena.
From Algorithmic Brilliance to Silicon Ambition
DeepSeek's meteoric rise was built on squeezing maximum performance from minimum compute. In January 2025, the R1 reasoning model stunned Silicon Valley by achieving breakthrough reasoning capabilities through reinforcement learning, at a fraction of the training cost of comparable models. Subsequent releases — V3, V4, and their variants — maintained the same philosophy: more performance with less compute.
Yet even the most brilliant algorithmic optimization ultimately hits physical ceilings. Founder Liang Wenfeng acknowledged in a rare 2024 media interview that chip shortages posed a real challenge. The R1 model was trained on NVIDIA H800 GPUs; thereafter, DeepSeek pivoted to Huawei's Ascend processors. The V4 model, released in April 2026, was adapted for Ascend, with Huawei confirming its processors participated in partial training of V4-Flash.
But placing chips in two baskets — NVIDIA and Huawei — was apparently not enough. DeepSeek's custom chip initiative reveals a deeper strategic calculus: in an era of mass AI commercialization, sovereignty over inference compute will define competitive moats.
Inference Chips: The Bottleneck of AI Commercialization
DeepSeek's choice to focus on inference rather than training is a shrewd strategic move. Training is a one-time investment; inference is a perpetual cost. Every user query, every API call, every deployed agent consumes inference compute. As AI applications scale from millions to billions of users, inference demand curves will rise exponentially.
Purpose-built inference chips can achieve significantly lower power consumption and per-unit inference cost compared to general-purpose GPUs. For a company racing toward large-scale commercialization, this is not merely a performance question — it directly determines whether the business model is viable. Estimates suggest that once API call volumes cross a critical threshold, the total cost of ownership (TCO) of custom inference silicon falls well below that of continued third-party GPU procurement.
Notably, DeepSeek had already laid the groundwork for hardware co-design at the model level. The UE8M0 FP8 data format introduced in V3.1 is widely believed to have been designed specifically for the hardware characteristics of next-generation domestic chips — the algorithm team was thinking about silicon while writing model code.
Why AI Companies Are All Building Their Own Chips
DeepSeek is far from alone. Across the globe, AI model companies are rushing toward chip independence.
In June 2026, OpenAI announced Jalapeño, its first custom inference chip developed in partnership with Broadcom, targeting inference workloads for the ChatGPT ecosystem. Anthropic was reported in April 2026 to be evaluating custom AI chip development to reduce reliance on third-party hardware. Further back, Google's TPU family has been deployed at massive internal scale, powering the training and inference of flagship models like Gemini.
Multiple forces are driving this trend. First, NVIDIA GPU supply constraints and premium pricing have become a universal pain point for any company deploying AI at scale. Second, general-purpose GPUs must accommodate graphics rendering, scientific computing, and diverse workloads, whereas inference-specific chips can shed redundant modules and achieve several times the energy efficiency on targeted workloads. Third, owning the chip enables deeper system-level optimization — from data center power delivery and thermal management to network topology and model architecture, everything can be co-designed around a custom silicon foundation.
Where the 51 Billion RMB Goes
Backing DeepSeek's chip ambitions is serious money. In June 2026, the company — which had steadfastly refused external investment for years — completed its first funding round, raising approximately 51 billion RMB (about 7.4 billion USD), with a post-money valuation between 52 and 59 billion USD.
The stated use of funds: expanding computing centers built on domestic chips, developing proprietary AI chips, and recruiting top global talent. On the infrastructure front, DeepSeek has posted job listings for IDC design and planning engineers, targeting data center projects from megawatt to gigawatt scale, with Inner Mongolia's Ulanqab explicitly mentioned among planned construction sites.
The chip design recruitment itself has been conducted with extreme discretion — no public job postings on any platform. This operational secrecy aligns with DeepSeek's signature approach: don't let the outside world know what you're doing until you have something to show.
Challenges and Outlook
Designing a competitive AI chip is no trivial endeavor. It typically requires years of development and billions of dollars in investment. DeepSeek has zero track record in this domain, and faces formidable barriers in chip architecture design, advanced-node tape-outs, and software toolchain ecosystem development.
But the flip side is compelling: DeepSeek possesses world-class model teams whose depth of understanding of AI workloads is something traditional chip design houses cannot match. The "model first, chip later" path ensures that silicon architecture precisely matches real inference demands, avoiding costly detours.
From a macro perspective, DeepSeek's chip initiative marks the entry of global AI competition into a new phase. The contest is no longer confined to model parameters and leaderboard rankings — it now extends upstream into the foundational infrastructure of compute. As major model companies pursue chip independence, "software-defined hardware" may be emerging as a new industry creed.
If DeepSeek's chip effort succeeds, it will not only reshape the company's position in the compute supply chain, but could blaze a trail for the entire Chinese AI industry toward "algorithm-to-silicon" vertical integration.
Sources: Reuters (2026-07-07), QbitAI (2026-07-08)