
DeepSeek, the Chinese AI startup behind some of the most downloaded open-source models of the past two years, is building its own AI inference chip. Three people familiar with the matter confirmed the effort to Reuters on July 7, 2026. The chip is designed specifically for inference, the phase where a trained model generates responses for real users, rather than for training new models.
Why DeepSeek Wants Its Own Silicon
Currently, DeepSeek depends on both Nvidia and Huawei hardware to train and run its globally popular models. A custom inference chip would give the company greater control over the systems powering its AI services while reducing exposure to supply chain restrictions.
DeepSeek has already shown willingness to work with Chinese-designed silicon. In April 2026, the company released its V4 model adapted for Huawei’s Ascend chips, and Huawei confirmed its processors were used in part of the training for V4-Flash, a lighter version of the model.
The Global Race for AI Inference Hardware
DeepSeek’s move mirrors a broader trend among AI companies seeking to reduce dependence on Nvidia. OpenAI announced in early July that it is co-developing a custom LLM inference chip codenamed Jalapeño with Broadcom. Google continues expanding its TPU ecosystem, and Amazon’s Trainium chips are gaining traction inside AWS.
Inference is where the real cost lives. Training a frontier model is expensive, but serving millions of daily users generates ongoing compute demand that dominates operating budgets. A purpose-built inference chip can deliver significant efficiency gains over general-purpose GPUs, processing more tokens per watt at lower latency.
What This Means for the AI Chip Market
Nvidia still commands roughly 80-90% of the AI accelerator market. But the company now faces pressure from multiple directions: cloud providers building custom silicon, AI labs designing their own chips, and Chinese firms like Huawei and DeepSeek working to bypass US export controls entirely.
DeepSeek’s inference chip is reportedly in early development stages, meaning a production timeline remains unclear. The company would likely face challenges around chip fabrication, since advanced semiconductor manufacturing requires access to TSMC or Samsung foundries, both of which face their own geopolitical constraints when dealing with Chinese entities.
For DeepSeek, the strategic logic is straightforward. The company’s models power millions of queries daily, and each one runs on rented Nvidia or Huawei silicon. Owning the inference stack from model to chip could dramatically lower serving costs while insulating the business from hardware shortages and export policy shifts.
Frequently Asked Questions
Is DeepSeek really building its own AI chip?
Yes. According to three sources familiar with the matter, reported by Reuters on July 7, 2026, DeepSeek is developing a custom AI inference chip to reduce reliance on Nvidia and Huawei hardware.
What is the difference between AI training and inference?
Training is the process of teaching a model using large datasets and compute resources. Inference is when the trained model generates outputs or responses for users in real time. Inference represents the ongoing cost of running AI models at scale.
How does this affect Nvidia?
Nvidia dominates the AI chip market, but faces growing competition from custom silicon projects by OpenAI, Google, Amazon, and now DeepSeek. The trend toward inference-specific chips could erode Nvidia’s market share over time.
When will DeepSeek’s chip be available?
The chip is in early development. No production timeline has been disclosed, and significant fabrication challenges remain before any commercial rollout.
