07/20 2026
525
When Claude demonstrated its astonishing ability to understand long texts, the computing power consumption behind it was equally staggering. It is estimated that the daily inference costs for maintaining Claude and other cutting-edge models with tens of millions of active users could reach hundreds of thousands of dollars per day in chip depreciation alone. This rate of burning money has forced Anthropic, a star company founded by former OpenAI core employees with a mission of AI safety, to seek a more fundamental solution beyond just financing. The answer points to self-developed AI chips, and its chosen ally is Samsung, which is eager to overtake in the AI foundry sector. This partnership has sounded the charge for large model companies to collectively move from the cloud to the silicon level.
01 Anthropic's Chip Ambitions
According to foreign media reports, Anthropic has entered into in-depth negotiations with Samsung Electronics, with the core topic being the development of an AI inference chip designed specifically for its large models. This is not a simple commercial procurement but a comprehensive technical collaboration covering architecture definition, manufacturing processes, and high-bandwidth memory integration. The news has shaken the industry because it means that large model companies' anxiety about computing power has extended from software-level "prompt optimization" and "model distillation" to the physical level of transistor circuit design.
The most direct driving force behind Anthropic's decision to bypass traditional chip vendors' off-the-shelf solutions and pursue self-development is the structural pain point of inference costs. General-purpose GPUs have a large amount of circuitry and functionality that are unnecessary for AI tasks when processing massive matrix operations, and their energy efficiency is far from optimal. A person close to the deal revealed that Anthropic envisions stripping away all unnecessary graphics rendering modules and building a dedicated ASIC chip that can improve efficiency by several times during specific model inferences, centered around its self-developed Claude model's sparse activation mechanism. This would not only significantly reduce the cost per API call but also free it from dependence on NVIDIA's CUDA ecosystem, truly putting the hardware in its own hands.
Choosing Samsung is a shrewd mutual pursuit. Samsung possesses integrated manufacturing capabilities ranging from chip design and wafer foundry to HBM high-bandwidth memory, which is extremely attractive for AI inference chips that are extremely hungry for memory bandwidth. For Samsung, winning Anthropic, a cutting-edge model company, as a customer means that its advanced processes below 3 nanometers and HBM memory have found an ideal partner that can continuously propose extreme demands and jointly define next-generation products. This is a crucial opportunity for it to rise in the foundry sector under the shadow of TSMC.
02 A Reverse Definition of the Industrial Chain
When Anthropic's actions are viewed in a broader context, it becomes clear that this is not an isolated case but a collective awakening sweeping through the entire circle of leading large model companies.

OpenAI has taken the most aggressive approach. Its CEO, Sam Altman, has been reported to be raising huge amounts of funds globally for a chip project codenamed "Tigris," with the goal of building a dedicated chip manufacturing network capable of supporting superintelligence. At the same time, OpenAI has collaborated with chip design giants like Broadcom to develop customized AI inference chips and has been in frequent contact with TSMC.
If OpenAI's aggressiveness is somewhat understandable—after all, it stands at the forefront of the large model race and feels the pain of computing power costs and supply security most directly—then when the list of companies developing their own chips continues to grow, the nature of the situation changes. Google's TPU has long been a textbook case of customized chips, with iterations up to the fifth generation, providing a powerful dedicated computing power base for its Gemini model. Microsoft's Maia 100 chip, released at the end of last year, directly targets cloud AI workloads, aiming to break NVIDIA's dominance in GPU procurement. Meta's journey in self-developed chips has been turbulent, but it has never given up, continuously developing more energy-efficient MTIA series chips for recommendation systems and generative AI. Even Amazon, which was once rumored to have abandoned self-development, has already deployed its Trainium and Inferentia chips on a large scale. From startups to cloud giants, from search dominators to social empires, almost every heavyweight player in the AI track (AI sector) has carved out its own territory in the silicon world.
Why are these companies, with vastly different business models and complex competitive relationships, do or think the same without prior consulation ( do or think the same without prior consulation , meaning "all moving in the same direction without prior agreement") heading down the same path? The answer may not lie in the individual strategies of each company but in the common computing power dilemma they face—a dilemma with three layers that progressively push all players toward the narrow door of self-developed chips.
The most superficial layer is the cost calculation. A cutting-edge large model can burn hundreds of thousands of dollars per day in chip depreciation alone during the inference stage. When the number of daily active users of a model exceeds tens of millions, using general-purpose GPUs means paying for a large amount of circuitry and functionality that are fundamentally unused in AI tasks. Self-developed chips eliminate all this redundancy, and even if energy efficiency improves by only 20-30%, it translates into a figure that can significantly change the financial report when calculated on an annual basis. According to the current price curve, the unit inference cost of a self-developed ASIC is expected to be one-fourth to one-third of that of a GPU of the same generation. This is not an optimization that adds icing on the cake but a life-and-death line that determines whether the model service can achieve a viable business closed loop ( closed loop , meaning "closed loop").
A deeper layer is supply anxiety. NVIDIA's GPU delivery cycles often stretch to more than half a year, with product rhythms completely beyond the control of downstream customers. All large model companies are aware of a cold reality: entrusting their technological lifeblood to a single supplier that simultaneously serves all their competitors is strategically unacceptable. Self-developed chips may not necessarily outperform NVIDIA in terms of performance, but they at least provide a trump card—when external supply issues arise, you still have your own production line to fall back on. In a sense, the insurance premium attribute of self-developed chips is more important than their technological value.
The most fundamental and decisive driving force is the deep integration of algorithms and hardware. In the past, model developers could only passively adapt to the architectural constraints of off-the-shelf chips; now, leading players are starting to do the opposite, tailoring chips to their models. Building models with a chip mindset and defining chips with model requirements—this soft-hard integration capability is becoming a new moat in AI competition. Google's success with TPU has already proven the power of this path: the inference efficiency of the Gemini model on TPU is far superior to that when transplanted to general-purpose GPUs. When every step of the algorithm's matrix operations can find precisely matching transistor circuits, performance improvements are not linear but leapfrogging. All large model companies want to replicate this leap.
These three pressures superposition ( superposition , meaning " superposition ," or "layered on top of each other"), have transformed self-developed chips from an option into a must for leading players. It's not that every company will succeed, but the cost of not trying has become too high to bear.
The collective chip development by large model companies is triggering a power restructuring in the chip manufacturing sector. Traditional chip design giants like NVIDIA and AMD need to re-examine the fact that their former customers are becoming potential competitors. However, for wafer foundries and design service providers, a feast has just begun.
TSMC is undoubtedly the biggest winner. Its advanced processes and CoWoS packaging have become virtually inevitable choices for all AI self-developed chips, with orders already booked for years to come. Samsung, on the other hand, is trying to seize market share from TSMC with its integrated memory-foundry-packaging bundle, and Anthropic's engagement is a potential sign of breakthrough for this strategy. Intel is rolling out its foundry service IFS and open chip interconnect standards, attempting to attract players who want to avoid TSMC's dominance. Meanwhile, the valuations of custom chip design service providers like Broadcom and Marvell are soaring as they develop multiple AI ASICs simultaneously for multiple giants, enjoying the dual benefits of selling shovels to those digging for gold.
This "chip development" movement is not a smooth path either. Designing a chip that can rival NVIDIA in terms of software ecosystem requires hundreds of millions of dollars in investment, with cycles lasting two to three years and filled with risks of failure. NVIDIA has erected a solid fence composed of hardware iteration speed, CUDA ecosystem, and NVLink interconnect technology. Jensen Huang recently publicly stated that even if competitors' chips were free, they might not be cheaper than NVIDIA's comprehensive cost. His confidence lies in the fact that NVIDIA is doubling performance every two years with a brutal evolution rhythm that leaves any self-developed chip vulnerable to obsolescence as soon as it lands. Self-developed chips are more like a gamble, using today's uncertainty to hedge against the even greater risk of being constrained tomorrow.

The partnership between Anthropic and Samsung marks a milestone in the extension of large model competition from parameter scale and multimodal capabilities to computing power sovereignty. When even the purest model companies start to define chips themselves, the vertical integration of the AI industry has become irreversible.
This wave is not exclusive to Silicon Valley. Across the ocean, Chinese AI companies are also accelerating their chip self-development efforts. Baidu has taken a full-stack self-development approach, with its Kunlunxin series AI chips already iterated to the third generation, featuring 7-nanometer process technology and tens of thousands of chips in mass production. These chips not only serve the training and inference of Baidu's ERNIE model but also have been scaled up in smart transportation, industrial internet, and other scenarios. ByteDance's moves are more secretive but equally determined, with a chip team of several hundred people focused on self-developed AI inference chips and server-specific chips, aiming directly at optimizing the inference costs of its Douyin recommendation system and Doubao large model. Huawei's Ascend series has shouldered the burden of domestic AI computing power under external sanction pressures, and according to public information, the Ascend 910C has demonstrated competitiveness against international mainstream products in multiple large model inference benchmark tests. Alibaba's Pingtouge's Hanguang series chips and Tencent-invested Suiyuan Technology are also continuing to invest in their respective paths.
The chip self-development practices of these Chinese players share the same underlying logic as those of OpenAI, Anthropic, and Google—all are trying to break free from dependence on a single supplier, all are using the reverse integration logic of defining hardware with algorithms, and all are stockpiling computing power sovereignty for the next stage of model competition. The only difference lies in the path: U.S. companies mostly take the route of joint customization with mature partners like Broadcom and TSMC, while Chinese companies, under technological blockades, are forced to take a more full-stack and autonomous breakthrough path.
In the future, the winners of the large model era may no longer win solely based on elegant algorithms but must also become hardware definers who deeply understand transistor physics and foundry games. Chip development has become a necessity rather than an option for leading AI companies in both China and the United States, a path to independence and long-term survival. When the boundary between algorithms and silicon completely disappears, this "soft-hard integration" war has only just begun.