09/22 2026
504
Imagine a machine no larger than a set-top box sitting on your office desk, equipped with hundreds of GB of unified memory, capable of running large models with hundreds of billions of parameters even without an internet connection. No server room, no racks—just plug it in and go, with data never leaving the room. This is the AI Box, a data center brought to the office desk.
In the AI Box arena, Apple, NVIDIA, AMD, Microsoft, and Intel are all present. But this form factor is not new. Why, in 2026, has this small box that 'becomes a computer when connected to a monitor' become such a hotly contested space?
There are two ways to use large models: via the cloud or locally.
Cloud-based AI offers unlimited capabilities, with no need to manage models yourself—the strongest capabilities are always available. But the costs are clear: under subscription models, services like ChatGPT Pro and Claude Max cap at $200 per month; beyond that, usage-based billing applies. For example, GPT-6 Astra, released on September 3, charges $50 per million tokens for output. Heavy users running agents that output tens of thousands of tokens per long task can easily burn through hundreds to thousands of dollars a month.
Industry insiders note that over the past two years, most enterprise AI adoption has been stuck in the exploratory phase of 'moving from concept to implementation': testing scenarios, overhauling processes, and fine-tuning prompts. Every minute in this phase costs money in tokens, yet bosses often provide only direction, not specific requirements. 'A lot of money has been burned, but what results have been achieved? No one knows.'
This is where the logic of local deployment comes in: most mainstream models are open-source and free (DeepSeek and Qwen series are both open-sourced under permissive licenses), hardware is a one-time purchase with no ongoing token bills, and data remains within the server room. Pre-research, validation, and repeated trial-and-error incur fixed costs.
Thus, the AI Box was born.
01 What qualifies as an AI Box?
Local AI deployment faces a hard technical barrier: the model must 'fit.' Memory is the entry ticket for an AI Box.
Running a large model requires loading all its weights into memory or video RAM (VRAM). The full-precision (FP16) version of DeepSeek-R1, with 671 billion parameters, requires approximately 1.3TB of VRAM; FP8 quantization reduces this to about 0.7–0.8TB, and even 4-bit quantization still needs 400–480GB. For comparison, an RTX 5090 has 32GB of VRAM, and a high-end laptop has 24–32GB of unified memory—only enough for small models with around 7 billion parameters.
This presents the biggest technical challenge for desktop AI devices: cramming as much unified memory as possible into a machine that fits on a desk, feeding the model the highest possible bandwidth, and keeping power consumption within what an office outlet can handle.
Memory capacity determines the size of the models that can run, bandwidth determines how many users can access it simultaneously and how fast text is generated, and power consumption determines whether it belongs in an office or a server room. The most dramatic twist in these three dimensions is that the company best prepared for this was not originally designing for AI.
In 2020, Apple transitioned Macs to its in-house M-series chips and made a then-radical decision: the CPU, GPU, and Neural Engine would share a single pool of memory—unified memory (UMA). This architecture, initially designed to eliminate the overhead of data transfers between chips and reduce power consumption, perfectly aligned with the demands of the large model era five years later.
Today's Mac Studio, in its M3 Ultra configuration, offers up to 512GB of unified memory, 819GB/s of memory bandwidth, and a maximum sustained power consumption of 480W. The M5 Ultra version, released on September 22, further increases bandwidth to 1.2TB/s and starts at 46,999 RMB in China. Apple officially states that this machine can run models with over 600 billion parameters. According to reports, OpenAI has purchased tens of thousands of Mac minis and Mac Studios over the past few months for reinforcement learning and agent development, with demand so high that Apple has been repeatedly urged to expedite shipments. Tim Cook acknowledged on July's earnings call that 'an increasing number of customers are using Mac mini as a powerful platform for running agents and deploying Mac Studio clusters for local operation of cutting-edge models,' noting that Mac supply has been constrained all season.
02 NVIDIA: Not selling machines, but 'the same architecture'
If Apple's approach was accidental, NVIDIA's is a deliberate 'dimensionality strike' after recognizing market demand.
The DGX Spark uses the GB10 Grace Blackwell superchip, featuring 128GB of unified memory, 273GB/s of bandwidth, and 1 PFLOPS of computing power at FP4 precision, with a total power consumption of 240W. NVIDIA claims it can infer models with up to 200 billion parameters as a standalone unit and 405 billion parameters when two are interconnected. Priced at $4,699 in the U.S. and around 30,000 RMB for the initial China release (now approaching 40,000 RMB due to memory price hikes), it offers no clear advantage over the Mac Studio on paper. However, the DGX Spark runs DGX OS (Ubuntu), sharing the same software stack as NVIDIA's data center superclusters, including CUDA ecosystem support and enterprise-grade tools. Developers can seamlessly transition work from their desktop to the cloud.
Around NVIDIA's DGX Spark, seven OEMs—Acer, ASUS, Dell, Gigabyte, HP, Lenovo, and MSI—simultaneously launched systems using the same chip. MSI's version starts at $2,999, and Lenovo's ThinkStation PGX entered China's government and enterprise procurement catalog in July last year. Compared to Apple, NVIDIA has a clearer market positioning and sales ambition for the AI Box.
If the Spark is an entry-level machine for developers, the DGX Station is a deskside workstation for small teams: featuring the GB300 chip, 748GB of coherent memory (252GB HBM3e + 496GB LPDDR5X), 20 PFLOPS at FP4, and 1,600W total power. NVIDIA claims it can run models with up to one trillion parameters. While its power consumption approaches the limit of standard office circuits, it remains an order of magnitude lower than data center servers.
At Microsoft's Build conference in June, the Surface RTX Spark Dev Box was unveiled: featuring the same GB10 chip as the DGX Spark, 128GB of unified memory, 1 PFLOPS of computing power, and a thermal design power of just 100W. Pre-installed with Windows 11 Pro for Developers, it claims to run models with up to 120 billion parameters and handle million-token contexts locally. Set for release in the U.S. in the second half of the year, media estimates place its price at $3,000–$3,500. Its significance lies more in its entry into the market than its specifications: the Windows ecosystem needed an official answer to show developers that local AI doesn't require buying a Mac.
03 The x86 camp: Demonstrating cost-effectiveness in the inference era
AMD's Ryzen AI Max+ 395 (codenamed Strix Halo) packs 128GB of unified memory and 256GB/s of bandwidth into a 55–120W power envelope, with an NPU delivering 50 TOPS and a total platform computing power of approximately 126 TOPS. The upcoming Ryzen AI Max 400 (Gorgon Halo) in Q3 will increase unified memory to 192GB and claim support for models with up to 300 billion parameters.
x86's strength lies in compatibility: Windows and Linux install directly, allowing enterprises to reuse existing operational frameworks. While not the highest-performing option, avoiding ecosystem changes is a major factor in procurement decisions.
Intel introduced Panther Lake (Core Ultra 300 series), featuring 50 TOPS of NPU computing power, integrated into every laptop and mini PC. Its closest box-like offering, the NUC 16 Pro, supports up to 96GB of memory and is positioned as a Copilot+ mini PC. Notably, Intel's AI Box strategy extends beyond desktop boxes to intelligent vehicles, enhancing computing power in previously underpowered cars. This reflects Intel's focus on specific B2B scenarios for AI Box deployment.
04 Who is actually buying AI Boxes?
What are these boxes being used for after purchase?
In higher vocational education, schools have developed three sets of agents for teachers, students, and administrators—covering lesson preparation, grading, practical training Q&A, and teaching plan refinement—all running locally on AI Boxes. Post-trial data shows: teachers' weekly lesson preparation time dropped from 9.2 to 5.1 hours, grading time per class fell from 2 to 0.4 hours, practical training error rates decreased from 22% to 14.3%, and the reuse rate of excellent teaching plans rose from under 10% to 67%.
Another scenario is law firms, where case files cannot leave the premises—ideal for AI Boxes. Case law retrieval times were reduced to under 40 seconds, automatic generation of legal documents covered over 80% of needs, and conflict-of-interest review accuracy exceeded 95%.
The logic behind these scenarios is clear: repetitive data processing + high data security requirements + verifiable results. AI Boxes shine here because traditional local server solutions require three to six months for server room modifications, while public cloud APIs offer instant access but risk data leakage. In contrast, AI Boxes offer plug-and-play deployment with physical isolation, measured in hours rather than months.
The shift from selling computing power to selling efficiency indicates that the market has moved beyond 'paying for parameters' and now prioritizes outcomes.
05 Conclusion
The potential of AI Boxes ultimately comes down to economics.
Building an in-house server room with eight RTX 4090 systems costs around 350,000 RMB on wholesale platforms, with peak power consumption exceeding 4,800W. This requires a dedicated server room, cooling, and engineer maintenance.
Cloud-based AI offers unlimited capabilities, but token bills grow linearly with usage, with exploration-phase costs potentially unlimited.
AI Boxes range from thousands to tens of thousands of RMB, consume 100–480W of power, and offer plug-and-play operation with free models and fixed trial-and-error costs. Compared to eight-GPU servers, power consumption is 10–30 times lower.
The industry still faces unresolved ROI calculations. Procurement decision-makers must re-evaluate based on their specific workflows. Scenarios like customer service, where AI can reduce headcount, are easy to quantify. However, process optimizations and efficiency gains are harder to measure. Additionally, like liability issues in autonomous driving, questions remain about who checks and takes responsibility for AI errors in practical work settings.
Currently, only two mature scenarios exist: coding and office work. Most other use cases are still exploratory. Hardware is merely the entry ticket: data governance, fine-tuning, and iteration occur in three-to-six-month cycles, and model changes often require reworking upper-layer applications. Another consideration is that boxes typically cannot be upgraded—capacity expansion requires purchasing additional machines. Selling boxes amid high DRAM costs may serve as a way to absorb expensive storage. The market needs clearer demonstrations of what AI Boxes can actually achieve.
AI Boxes are not here to replace the cloud. Tasks requiring the strongest models, elastic computing power, and zero maintenance still belong in the cloud. AI Boxes address a different need: exploratory pre-research with repeated trial-and-error, data that cannot go to the cloud, fixed computing power for small teams, and cost-effective operations for small businesses.
By miniaturizing hardware, AI Boxes enable computing power freedom after a fixed upfront cost. Ultimately, the product's future depends on how computing power is implemented to deliver real value.