10/08 2026
357
In recent years, Silicon Valley has increasingly resembled a sprawling construction zone. Microsoft, Google, and Amazon are continuously expanding their data centers, with competition extending beyond GPUs to encompass electricity, land, and cooling systems. Tech companies that once relied on software and network effects for rapid growth now face soaring capital expenditures.
As the data center race intensifies, both Musk and Microsoft are shifting their focus to opposite ends of the computing power supply chain.
Musk has confirmed that Tesla and SpaceX will construct and operate their own Terafab chip manufacturing facilities, with TSMC potentially leasing some of the space. Meanwhile, Microsoft, in collaboration with NVIDIA, has introduced the Surface Laptop Ultra, featuring RTX Spark chips, and unveiled its Windows Hybrid Intelligence solution. This enables certain large AI models and agent tasks to run directly on personal computers.
One company is preparing to venture into semiconductor manufacturing, while the other is attempting to shift more computing tasks back to personal devices. Despite facing different technical challenges, both are addressing the growing contradictions within the AI industry.

As AI evolves from an occasional tool into a service that continuously consumes computing resources, the methods for supplying computing power and the associated costs have become more complex than simply scaling up infrastructure.
01 Why Is Musk Determined to Build His Own Chips?
When Morris Chang founded TSMC in 1987, the semiconductor industry was dominated by vertically integrated companies like Intel and Texas Instruments, where chip design and manufacturing were largely conducted in-house. TSMC specialized in manufacturing chips for customers, gradually transforming the industry’s division of labor and allowing companies like NVIDIA and AMD to bypass the enormous investment required to build their own fabs.
This division of labor has a solid economic foundation. By 2025, TSMC will employ 305 process technologies, manufacturing 12,682 different products for 534 customers. Cross-industry orders help spread equipment depreciation, balance capacity utilization, and continuous production accumulates experience to improve yields. Advanced processes, in particular, rely on such accumulation: even with the same equipment, different fabs can vary significantly in yield and unit costs.
Apple, with its world-class chip design team, still primarily entrusts advanced process chips to TSMC for manufacturing, exemplifying this industrial division of labor.
However, Musk now seeks to cross this boundary.
According to SpaceX's filing with the U.S. Securities and Exchange Commission, the Terafab project plans to integrate logic chip, memory manufacturing, and advanced packaging, with a long-term goal of producing hardware equivalent to 1 terawatt of computing power annually. Here, "terawatt" refers to power capacity, not computational speed, and certainly does not mean the factory has already achieved corresponding production capacity.
The project aims to serve Tesla's automotive and humanoid robot businesses, as well as SpaceX's planned orbital AI computing. For Musk, when future chip demand becomes large enough, the expansion speed, delivery schedules, and capacity allocation of external foundries could impact the launch rhythm of his products.

Musk's concerns about the supply chain also include geopolitical factors. At the September All-In Summit, Musk publicly expressed worries that chip supplies from Taiwan could be disrupted for various reasons in the future. Even excluding geopolitical risks, the expansion capabilities of existing fabs may not meet the long-term needs of AI, automotive, and robotics.
Economist Coase raised a question in his 1937 work, The Nature of the Firm: Why are production activities sometimes entrusted to the market and sometimes retained within a firm? A key explanation is transaction costs. When the costs of external procurement and coordination rise, firms may expand their own boundaries. Terafab can be understood in this context, but the particularity of semiconductor manufacturing is that internal organization is also expensive. A fab primarily serving self-use demands must alone bear the risks of yield improvement, idle capacity, and technological obsolescence.
Therefore, Musk's plan does not yet prove that vertical integration is superior to professional foundry services. SpaceX's disclosure explicitly states that it will still procure a significant proportion of computing hardware from third parties; while Intel has confirmed its continued participation, specific investment and construction arrangements remain to be determined.
Musk hopes to reduce dependencies on others but may also transfer manufacturing risks originally borne by suppliers onto himself.
02 Why Is Microsoft Bringing AI Back to Personal Computers?
Compared to Musk, Microsoft's moves also involve costs but in the opposite direction.
Azure has built a massive cloud business relying on centralized computing, yet Microsoft now wants Windows to handle more local inference tasks. This shift first occurs in the models themselves.
Microsoft's MAI Code 1.1 Flash local version boasts 137 billion total parameters but activates only about 6.8 billion parameters when generating each token. It adopts a Mixture of Experts (MoE) architecture, using a routing mechanism to call only part of the expert networks, reducing computation per step; however, inactive parameters still need storage, so quantization compresses model weight usage to about 53GB, an ~80% reduction from high-precision versions.
In Microsoft's 256K context test, the model's peak memory usage was 75.5GB; in the SWE-bench Verified test, the local quantized version scored 70.8%, close to the high-precision version's 72.6%. This suggests model compression can retain task capabilities to a certain extent, but these results belong to vendor-published specific tests, with actual speed and effectiveness still influenced by memory bandwidth, context length, and other conditions.
This explains the significance of the Surface Laptop Ultra's maximum configuration featuring 128GB unified memory. However, local large models have not suddenly become cheap: the product starts at $2,599, with the entry version offering only 24GB memory and the 128GB version nearing $5,900. High-end machines currently target developers and professional users primarily.
Microsoft is clearly vying for Apple's high-end developer market. During the launch event, Microsoft compared the RTX Spark platform's local AI performance with the 16-inch MacBook Pro equipped with the M5 Pro chip and offered a trade-in rebate of up to $1,000 for MacBook Pro users. For Microsoft, making Windows the preferred platform for high-performance local AI development again is equally an important goal of this hardware upgrade.
Microsoft is betting on changes in usage habits. In the past, people occasionally asked chatbots questions, making cloud computing convenient; today, code agents may continuously read files, run tests, and modify programs, requiring multiple model calls per task. Offloading simpler, repetitive tasks to local devices could reduce remote call costs and help control latency and sensitive data transmission.
Windows Hybrid Intelligence thus gains commercial significance: complex tasks can still call the cloud, while suitable local tasks are handled by user devices.

GitHub Copilot's local and cloud scheduling capabilities are planned for experimental preview later in October, with actual cost advantages yet to be verified. After all, purchasing hardware is also an expense, and when utilization is low, cloud services may be more cost-effective. However, once agents become resident applications on personal computers, Microsoft has the opportunity to extend Windows' OS advantages into the AI era, further controlling the scheduling gateway for local and cloud computing.
03 Computing Power Competition Begins to Redefine Cost Boundaries
The computer industry has experienced migrations from mainframes to personal computers and then to cloud computing. Each concentration and decentralization of resources has been influenced by device prices, network capabilities, and usage demands. AI is driving new adjustments but will not simply replay the history of personal computers replacing mainframes: frontier model training and high-intensity inference still require large clusters, while increased local applications may, in turn, stimulate cloud demand.

Terafab attempts to bring more capital and manufacturing risks in-house in exchange for supply control; Windows Hybrid Intelligence disperses some computing to user devices, seeking more flexible task costs. The former reevaluates the costs of external foundries, while the latter reevaluates the costs of cloud calls. Both point to a question: How much will it ultimately cost an enterprise to acquire and use one unit of effective AI capability?
This question is changing how AI competitiveness is measured. GPU counts, model parameters, and data center areas can reflect investment scale but struggle to directly indicate economic returns. Large clusters with insufficient utilization will see depreciation eat into profits; smaller models that complete tasks at lower costs may possess stronger commercial value.
Of course, TSMC's scale advantages remain firm, and cloud computing is far from losing efficiency. Whether Musk can establish advanced manufacturing capabilities and whether Microsoft can save users enough money through local computing will require real-world operational results to answer.
Data centers will continue to expand, but the competition for control over the AI industry has begun to extend to both ends of the supply chain. Musk attempts to incorporate chip manufacturing into his industrial system, while Microsoft hopes to leverage Windows to control the scheduling gateway for local and cloud computing. The two companies have adopted different strategies but both aim to secure more advantageous positions in the future production and distribution of computing power resources.
In recent years, tech giants have continuously proven their ability to build larger data centers. In the future, those who can influence how computing power is produced, where it runs, and who bears the costs will have a greater opportunity to seize greater industrial initiative.
After all, possessing computing power is one form of strength; deciding where computing power comes from and where it goes is another kind of power.