07/21 2026
410
When you open the price list of any AI cloud service provider, the most striking entry is invariably the cost for "input/output per million tokens." This billing model, which the industry humorously dubs the "AI utilities," has been the dominant paradigm for nearly three years. Developers watch their usage metrics spike on dashboards, their hearts racing, while self-mockingly referring to themselves as "digital-age water carriers" on forums—every keystroke triggers a mental calculation of cost.
However, over the past six months, a wave of unconventional AI products has quietly shifted their pricing metrics from "tokens" to "tasks completed" or even "effective deliverables." This subtle shift represents a far more profound transformation: a redefinition of how software value is measured.
01. Paying for Word Count or Paying for Value?
At first glance, token-based billing seems logical. The model processes inputs and generates outputs, with each step consuming GPU computational resources. Charging by tokens resembles paying a translator by the word—fair, transparent, and quantifiable.
But users quickly notice a disconnect: the model's "effort" rarely aligns with the user's "perceived value."
Consider this example: An operator of a cross-border e-commerce platform once shared with industry media that he requested a competitive analysis report using 2,000 tokens. The model returned a verbose 3,000-token document filled with correct but irrelevant phrases like "overall" and "it is worth noting," with only three to five lines of actionable insights. Yet he was billed for all 5,000 tokens—because the platform measured value in tokens, not usefulness.
This pricing model essentially treats the model's internal computational costs as the sole basis for pricing, shifting all the risk of value judgment onto the user. Economically, this reflects a classic "cost-plus" mentality—appropriate for the industrial age but ill-suited for AI, whose mission is to solve complex problems.
A less obvious cost lies in user psychology. When every interaction is instantly translated into token consumption, a form of "measurement anxiety" sets in. Product managers avoid using AI for brainstorming because divergent thinking wastes tokens; legal advisors hesitate to let the model compare contract clauses line by line because long texts blow budgets.
Users begin to behave like homeowners obsessively monitoring their water meters, wincing at the cost every time they turn on the tap. This psychological burden distorts AI's true usage scenarios—a tool designed for exploration and experimentation is reduced to a cheap typewriter used only for "quick and easy" tasks.
Research institutions have statistically observed that under token-based billing, over 60% of enterprise users deliberately shorten their queries and omit contextual details, sacrificing AI's strengths in long-term reasoning and holistic understanding. In other words, the meter itself alters driving behavior, preventing the vehicle from reaching its full potential.
02. The Weight of the Word "Outcome"
Amid this collective anxiety, the "outcome-based payment" model has begun to take root in niche markets.
Early adopters emerged in marketing copy generation. Several AI writing startups now charge not by characters but by "usable final drafts"—users submit requirements, the tool generates three versions, and the user selects one for refinement. The system bills only for the chosen draft, even tiering prices based on its subsequent click-through rate. Later, code assistance tools adopted a similar logic: charging not for lines of code written but for "PRs (Pull Requests) merged into the main branch and passing tests."
The number of intermediate versions generated by the AI or the debugging conversations discarded are absorbed by the provider. For users, the bill reflects a clean metric: how many problems were solved or how many effective reports were produced that month.
This shift appears to be a pricing adjustment but fundamentally reverses risk assumption. Under token-based billing, users were like farmers' market shoppers—they weighed and paid for produce, then took it home to wash, chop, and cook. Whether the dish tasted good or suited their taste depended on their culinary skills, with the vendor bearing no responsibility. Outcome-based payment, by contrast, resembles dining at a restaurant—customers describe their desired flavor and budget, while the kitchen manages ingredient waste and offcuts. If the dish is unsatisfactory, the restaurant must remake it or comp the meal.
AI providers transition from "selling computational power" to "selling deliverables." They must now truly understand the problems users need to solve rather than mechanically responding to queries. This forces vendors to invest more in intent recognition, multi-round clarification, and automatic error correction, as every ineffective output becomes their cost, not the user's.
Of course, this transition is not without challenges. The primary controversy centers on defining "outcome." What constitutes success? Does a copy's open rate matter? Does code runtime count? How should a consultancy report's adoption rate be quantified?
Standards vary wildly across scenarios, leading to potential disputes. Some enterprises have tried using "manual user confirmation" as a validation threshold, but this introduces new cognitive burdens and manipulation risks. A thornier issue arises technically: to reduce costs, might models output shorter, more conservative responses to end conversations quickly? If "outcome" is crudely defined as "answer length below a threshold," it could push AI toward mediocrity.
These controversies make clear that outcome-based payment is not a universal solution. It works best for tasks with clear boundaries and objective acceptance criteria, while the token model remains practical for creative, exploratory dialogues.
03. From Counter to Value Scale
Nevertheless, the shift from tokens to outcomes is reshaping the cost structure and competitive logic of the AI service industry.
Previously, cloud providers competed on computational power affordability or concurrency—essentially a resource race. Future entrants may differentiate themselves by achieving higher "outcome success rates" or lower "ineffective rounds." Some leading platforms have already adopted hybrid billing—charging a nominal per-token fee for basic dialogues while pricing complex tasks (e.g., financial report analysis, legal research) by deliverables. This resembles how power companies now bundle services based on "lighting duration" or "cooling effectiveness," which better align with user needs.
From a longer-term perspective, token-based payment was a natural product of AI's early industrialization—when model capabilities varied widely, computational resources were scarce, and consumption-based pricing was the fairest metric. Today, however, model capabilities are rapidly reaching a baseline of "passable utility." What truly differentiates them is no longer text volume but the depth of intent understanding and solution precision.
As software evolves from a "tool" into a "digital employee," users naturally refuse to pay for idle time—they want to pay only for results. This measurement revolution may be closer to commerce's essence than the arms race of model parameters.
- The End -