跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

In Four Years, Could AI Token Calls Be 10,000x Cheaper?

Original · Unique Research · 2026-07-15

Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening PPIO demo anecdote, six insight sections, and the full 18-question Q&A transcript. Cost, price, performance, and usage figures are source or speaker claims attributed to Yao Xin / PPIO, not independently verified findings. Product names, company names, and named people are preserved as source attributions; official English forms are used where known and transliterated where unverified.

AI Industry Observation

In Four Years, Could AI Token Calls Be 10,000x Cheaper?

"The truly important infrastructure of the AI era is not piling up more GPUs, but making every single token call smarter and cheaper."

PPIO once demonstrated a task that wasn't especially complex: give the AI no preset conditions and let it figure out how to operate the computer on its own. When the task finished, the bill was only a few yuan — but behind the scenes, nearly 500,000 tokens had quietly been consumed.

Many AI application entrepreneurs have probably never done this math. They worry more about whether the model works well and have no time to calculate how much each individual call actually costs. Yet this is precisely what PPIO co-founder and CEO Yao Xin (姚欣) cares about most right now. We sat down with him to talk about everything from token cost to Agent infrastructure, from the paradox of compute being "both scarce and surplus," to his judgment on the AI startup bubble.

The Question Has Changed: From "Can the Model Do It?" to "Can You Afford It?"

If you stripped away the official marketing language, what is PPIO actually doing? Yao Xin puts it this way: "PPIO is building a 'token factory' for Agent demand. We make every token call for AI entrepreneurs cheaper, faster, and more stable. For the coming Agent era, we are building an Agentic Cloud that lets AI developers run inference, operate, and deploy Agents at low cost."

When he looks at any emerging field, he watches two indicators: technology usage cost on the supply side, and penetration rate on the demand side. Take mobile internet — early iPhones were expensive, then Xiaomi launched thousand-yuan phones, and mobile data prices fell from tens of yuan per GB in the 3G era to a few jiao in the 4G era. An order-of-magnitude drop in cost was the key to mass adoption.

AI is following the same path, but at a much faster pace. Yao Xin did the math: an Agent task typically loops autonomously 10 to 50 times — think, call tools, observe results, think again — and each round must process the full context, with computational load growing larger every time. "This isn't linear growth; it's exponential ballooning." The industry's focus over the past two years has been "can the model do it," but as DeepSeek, Qwen, GLM and other models grow increasingly capable, the real question has become "can you afford it?"

"Token cost will drop 10,000x in four years. An inference that costs one cent today may cost a thousandth of a cent in the future."

But the supply-side pressure is real. The most expensive component of an AI GPU is HBM high-bandwidth memory, which accounts for 50% to 70% of material cost. HBM prices have nearly doubled over the past two years. Globally, only SK Hynix, Samsung, and Micron can manufacture HBM, and the 3D stacking process requires a 12- to 18-month expansion cycle amid supply shortages. From HBM price hikes, to GPU cost increases, to cloud vendor procurement cost rises, to inference service price adjustments — it is a clear transmission chain.

On one side, unit prices are falling. On the other, usage is exploding exponentially. "What PPIO does is help push the per-token price down." Intelligent model routing, semantic caching, and elastic inference scheduling are all fundamentally about lowering the startup cost and marginal cost of AI applications.

Compute Is "Both Scarce and Surplus": Where Do Entrepreneurs Most Often Trip Up?

Yao Xin summarized the current compute situation in one sentence: both scarce and surplus — essentially a structural mismatch, not a total-volume problem. What's scarce is effective compute equipped with the latest GPUs, high-speed interconnects, and complete software stacks. What's surplus is "bare metal" with outdated hardware selection, weak operations capability, and remote geographic locations. Many intelligent computing centers were originally built for training scenarios, but the market has shifted toward inference and Agents — inference requires elasticity, low latency, and on-demand invocation, which misaligns with the design intent of centralized data halls.

Centralized cloud vendors' inference pricing includes a large infrastructure premium, making per-token prices too high for AI applications already operating on thin margins. Moreover, AI application inference demand fluctuates wildly: an Agent can spike from zero calls to 1,000 concurrent requests in a single second, and traditional cloud scaling speed cannot keep up. Real-time AI interaction is extremely latency-sensitive; a centralized data center may be thousands of kilometers from the user, while edge inference nodes can compress latency to 20 to 50 milliseconds. This is why PPIO abandoned large-model training and focused on inference — strategy is about trade-offs.

On the specific pitfalls entrepreneurs most often fall into, Yao Xin listed several. First, calling the most expensive SOTA model directly during the Demo phase — results look great but costs can reach $10 to $30 per million tokens, when 80% of scenarios only need a mid-tier model at one-tenth the cost. Second, skipping prompt optimization and semantic caching — a well-crafted prompt can save half your tokens, and with semantic caching, many teams discover after launch that 30%+ of calls are actually duplicates. Third, provisioning inference resources for peak load when actual utilization is only 20% to 30%.

"Don't look for nails just because you have a hammer. Don't treat 'compute' itself as the goal. What entrepreneurs really need is 'a service that actually runs.' Between compute and service, there is still the software stack, operational capability, and user experience layer."

What Entrepreneurs Should Care About Is Cost Structure, Not Underlying Architecture

"AI application companies don't need to worry about the underlying compute architecture — that's why we exist," Yao Xin said. "But they must care about the inference cost structure: how much each call costs, and what gross margin looks like after scaling. 'Choosing the right infrastructure partner is essentially solving this business math for yourself.'"

The shift from Chat to Agent has completely changed the underlying infrastructure requirements. Yao Xin breaks it down into three layers. First, resource management has shifted from "short spikes" to "long duration" — generative AI completes an inference in a few seconds and releases the GPU quickly; agentic AI may run a task for minutes or even hours, continuously calling the model, so the same user volume can amplify compute demand by two to three orders of magnitude. Second, token cost optimization — automatically selecting the right model, automatically caching prompts, automatically routing to low-cost nodes. Third, execution environment security and compliance — AI-generated code is unpredictable and needs a "safety fence"; an Agent needs not just an inference API but a complete execution environment: file system, network access, snapshot recovery.

He also raised an interesting contradiction: AI has an "anti-organizational" character. Large language models hallucinate, which is essentially the uncertainty of generative AI. But what large institutions pursue in high-value scenarios is precisely certainty, which naturally conflicts with AI's character. One current industry approach is to run multiple Agents in parallel, trading inference consumption for reliability. "This complexity should be borne by the platform, not solved by every startup team on its own."

On PPIO's own capability stack, there are three layers. The model service layer is an intelligent token factory for the Agent era — platform-neutral, model-rich, with over one trillion daily token calls and 560,000+ registered developers. The Agent Harness platform layer is China's first Agent sandbox compatible with E2B, integrating multi-agent orchestration, Tool Use, and memory management; its business scale grew over 100x in the year since launch. The Agent access layer ships prebuilt templates for OpenClaw, Hermes Agent and others, supporting AI-native interfaces like MCP/Skills so Agents can directly call infrastructure resources.

"Entrepreneurs only need to care about two things: product logic and user experience."

What Customers Buy Isn't Compute — It's "A Service That Runs"

One innovative company building a one-stop AI assistant combining chat, search, collaboration, and coding support came to PPIO with two concrete problems: it needed to quickly validate new feature feasibility without wasting resources, and most LLM APIs on the market lacked consistency, reliability, and cost-effectiveness. PPIO's solution was optimized distributed inference and inference acceleration, plus a DeepSeek R1 API with function-calling support, helping the team rapidly test features in role-playing, story expansion, and code generation scenarios and accelerate development iteration.

Behind these cases is Yao Xin's judgment on the business model shift: "Customers buy exactly three things — model APIs, Agent sandboxes, and GPU services with enterprise private deployment." On the model API side, PPIO uses a unified interface to connect mainstream open-source models including DeepSeek, Qwen, GLM, Kimi, and MiniMax, covering large language models, image generation, speech recognition, and video generation, billed by actual token consumption. The Agent sandbox is runtime infrastructure built specifically for AI Agent scenarios — executing code and accessing file systems in isolated cloud environments, because "you can't let it run bare in your production environment." GPU services and private deployment target large customers with compliance requirements or dedicated resource needs.

The results are concrete: in 2026, PPIO's input token pricing for well-known large models delivered through its model API service is over 40% below the average rates of major international cloud providers. For the same model, compared to deploying with a baseline inference engine like SGLang, first-byte latency dropped 95% and total throughput efficiency improved 300%.

Yao Xin specifically emphasized something easily confused: PPIO is not a "GPU rental platform" nor a "model API aggregator." "Simply reselling compute is a business but hard to build into a long-term venture." What they provide is inference services covering the full lifecycle of AI applications — intelligent model routing, token cost optimization, Agent runtime environments, and elastic inference scheduling as a complete capability set. "Developers don't need 'compute'; they need 'a service that runs.'"

Their relationship with the industry chain is also clear: upstream/downstream with large model companies — "they build the engines, we do intelligent routing and inference optimization"; a service relationship with application companies; and more complementary than competitive with traditional cloud vendors, because big-cloud inference services are usually tied to their own ecosystems and models, while PPIO as an independent third party provides neutral routing not locked to any single cloud or model. On long-term moats, he summarized four: quality compute resource reserves, inference algorithm optimization capability, API compatibility and platform generality, and end-to-end Agent-facing capabilities — secure execution environments, state management, workflow orchestration.

"AI That Doesn't Talk About Revenue Implementation Is a Bubble"

Yao Xin has been an investor himself and ran an AI startup camp during a visiting scholar period in Silicon Valley. Returning to the entrepreneurial frontline, he admitted: "I may not be able to return to the state I was in at 24, so we choose to support today's 'Yao Xins' and accompany them to success."

On the current "heat" of AI entrepreneurship, his judgment is direct: "AI that doesn't talk about revenue implementation is a bubble." The directions he sees favorably are AI coding, AI Agents (enterprise automation), AI video/content generation, and AI education — scenarios with clear willingness to pay and quantifiable efficiency gains. He is especially high on AI coding because "code means automation and intelligence — if AI can generate code, it means it can control computers." PPIO itself started from pain-point scenarios with low accuracy requirements but willingness to pay, such as marketing, meetings, and back-office tools.

"For entrepreneurs, he gave three indicators. First, token cost per interaction — this is the 'raw material cost'; if you can't calculate it clearly, your business model is a castle in the air. Second, user penetration — borrowing from Crossing the Chasm, reaching around 20% penetration means the technology is moving from early adopters to the mainstream. Third, the LTV/CAC unit economics model — not simply customer acquisition cost, but lifetime value per paying customer divided by acquisition cost."

Next Year, Masses of Agents Will Take Over Existing Workflows

Over the next three years, Yao Xin expects per-token prices to drop another one to two orders of magnitude while usage grows four to five orders of magnitude. The shift from Chat to Agent will amplify compute demand for the same user base by two to three orders of magnitude — generative AI is "short spikes," agentic AI is "long duration" — not short-term peaks, but sustained high-concurrency demand. "Next year is expected to be the Year of the Agent, when masses of Agents begin taking over existing workflows." For the same spend, AI capability improves tenfold every year, a trend that will directly spawn a batch of "free AI applications."

"The truly important compute infrastructure of the AI era is not piling up more GPUs, but making every token call smarter and cheaper. This is consistent with the essence of technology development: continuously improving the efficiency of resource utilization."

Interview Highlights Q&A

Q1: If you stripped away the official introduction, how would you explain to someone unfamiliar with PPIO what you're doing?

Yao Xin: Focusing on the present, PPIO is building a "token factory" for Agent demand, making every token call for AI entrepreneurs cheaper, faster, and more stable. For the future Agent era, we are building an Agentic Cloud so AI developers can run, infer, and deploy Agents at low cost. We have chosen to be an empowering infrastructure company — focused on being the foundation that supports outstanding entrepreneurs as they grow.

Q2: Over the past two years, attention has been on model capabilities and application innovation. Why do you think compute infrastructure is increasingly becoming the core issue?

Yao Xin: When I look at any emerging field, I focus on two core indicators: technology usage cost on the supply side, and penetration rate on the demand side. Take mobile internet as an example — early iPhones were very expensive, then Xiaomi launched thousand-yuan phones, and data rates fell from tens of yuan per GB in the 3G era to a few jiao in the 4G era. An order-of-magnitude cost decline is key to mass adoption. AI is following the same path, but token consumption is growing far faster than traditional internet traffic. An Agent task may autonomously loop 10 to 50 times: think, call tools, observe results, think again. Each round must process the full context, and computational load grows larger each time. This isn't linear growth; it's exponential ballooning. Over the past two years, the industry asked "can the model do it?" But as DeepSeek, Qwen, GLM and other models grow increasingly capable, the real question has become "can you afford it?" Token cost is becoming the decisive variable for whether an AI application's business model works.

Q3: You've mentioned that AI applications may enter a "free era." If application prices keep falling, what pressure does that put on underlying inference costs?

Yao Xin: We predict that token cost will drop 10,000x within four years. An inference that costs $0.01 today may cost $0.001 or less in the future. But supply-side cost pressure is real. The core cost component of an AI GPU is HBM high-bandwidth memory, accounting for 50% to 70% of material cost. HBM prices have nearly doubled over the past two years, and only a handful of manufacturers globally can produce it, with expansion cycles typically 12 to 18 months. From HBM price hikes to GPU cost increases to cloud vendor procurement and inference service price rises, it's a clear transmission chain. What PPIO does is help push per-token prices further down. Intelligent model routing, semantic caching, and elastic inference scheduling all fundamentally reduce the startup and marginal costs of AI applications. One critical trend ahead: per-token prices keep falling while total token usage grows four to five orders of magnitude. Falling unit prices and exploding total volume — that is precisely the opportunity for infrastructure companies.

Q4: What cost traps do AI application entrepreneurs most often fall into?

Yao Xin: The first trap is using the largest model to solve every problem. Many entrepreneurs call SOTA models directly for Demos — results look great, but costs can reach $10 to $30 per million tokens. In reality, 80% of tasks only need a mid-tier model at one-tenth the cost. We once demonstrated AI operating a computer without preset conditions, and a simple task consumed nearly 500,000 tokens. It's technically feasible, but the cost is nowhere near low enough. The second trap is skipping prompt optimization, caching, and context management. The same task can save 50% of tokens with a well-crafted prompt. Combined with semantic caching, many duplicate questions can be skipped entirely. We've seen teams discover after launch that 30%+ of calls are actually duplicates. The problem is worse in Agent mode — every round processes the full context, and without context compression and tiered caching, costs spiral like a snowball. The third trap is provisioning inference resources for peak load when daily utilization is only 20% to 30%. The fourth trap is overindulging in technical optimization while ignoring the balance between technology and business — you can't look for nails just because you have a hammer. The fifth trap is treating "compute" itself as the goal. What developers really need isn't a GPU but a complete service they can call, an Agent they can run, enough elasticity, and a reasonable price.

Q5: Has compute cost already become the decisive variable for whether an AI application's business model works?

Yao Xin: Yes, and it already is — but it varies widely by scenario. For example, AI companion and social products have high conversation frequency, high per-interaction token consumption, and long contexts, but low ARPU, making them extremely sensitive to per-token prices. Our platform's customers at the beginning of last year were mainly chat and companion apps. This year, a wave of productivity tool applications has emerged. These products require multi-round model calls per task, and their unit economics are equally fragile. An interesting phenomenon is that our platform's top customer list reshuffles significantly every six months to a year. This shows the AI application ecosystem iterates very quickly. The teams that ultimately survive are often the ones that figure out token cost earliest.

Q6: When should an AI application company start thinking about compute and inference costs?

Yao Xin: AI application companies don't need to research underlying compute architecture themselves — that's why PPIO exists. But from day one, they must care about inference cost structure: how much each call costs, what gross margin looks like as user scale grows, and whether the product's unit economics work. Choosing the right infrastructure partner is essentially solving this business math for yourself.

Q7: Why can't existing cloud infrastructure fully meet AI application and Agent needs?

Yao Xin: Today's compute is "both scarce and surplus" — essentially a structural mismatch, not simple total-volume insufficiency. What's truly scarce is effective compute equipped with the latest GPUs, high-speed interconnects, and complete software stacks. What's in surplus is "bare metal" with outdated hardware selection, weak operations capability, or inability to adapt to actual application needs. Many intelligent computing centers were originally built for large-model training scenarios, but the market is shifting from training to inference and Agents. Training prioritizes large-scale, centralized computation; inference requires elasticity, low latency, and on-demand invocation. The needs are completely different. AI application request fluctuation is extreme — an Agent can spike from zero to thousands of concurrent requests in a single second. Traditional cloud scaling speed cannot fully match this demand. This is why we chose to abandon large-model training and focus on application inference. Strategy is about trade-offs, choosing based on external environment and your own capabilities.

Q8: What problem does one-stop AI cloud service solve most directly for startup teams?

Yao Xin: The most direct thing is "don't reinvent the wheel." An AI startup team typically has only 3 to 5 people. If they also have to deploy models, optimize inference, manage GPUs, and handle operations, half their energy is wasted on infrastructure. PPIO's capabilities are divided into three layers. The first is model service — an intelligent token factory for the Agent era, connecting DeepSeek, Qwen, GLM, Kimi, MiniMax and other models through a unified API, covering multimodal capabilities in language, image, voice, and video, focused on token cost and long-horizon task problems. The second is the Agent Harness platform, including an E2B-compatible Agent sandbox, multi-agent orchestration, Tool Use, and memory management, providing a complete runtime environment for Agents. The third is the Agent access layer, shipping prebuilt templates like OpenClaw and Hermes Agent, and through MCP, Skills and other AI-native interfaces, letting Agents directly call infrastructure resources and services. Entrepreneurs should focus their energy on two things: product logic and user experience.

Q9: From Chat to Agent, what fundamental changes happen in underlying infrastructure?

Yao Xin: From Chat to Agent, AI shifts from being "a mouth that answers questions" to "a person who does hands-on work." Doing work requires tools, environments, continuous thinking, and memory — which makes completely different demands on underlying infrastructure. First, resource management shifts from "short spikes" to "long duration." Generative AI inference may finish in a few seconds and release the GPU quickly; an Agent task may run for minutes or even hours, continuously calling the model. The scheduling system no longer just handles rapid response but long-task management and sustained resource occupation. For the same user volume, Agent-driven compute demand can amplify by two to three orders of magnitude. Second, the platform must automatically optimize token cost, including selecting the right model, caching prompts, compressing context, and automatically routing to lower-cost inference nodes. Third, Agents need a secure execution environment. AI-generated code is unpredictable and cannot run bare in production. It needs a secure sandbox with file system, network access, permission isolation, and snapshot recovery. This complexity shouldn't be rebuilt by every startup company; it should be uniformly borne by the infrastructure platform.

Q10: What is the relationship between model API-ification and inference service-ification?

Yao Xin: They are complementary, not substitutive. Model API-ification lets developers avoid deploying models themselves, but calling the API itself consumes tokens. When application scale is small, directly calling model API is usually sufficient. But as user scale grows, token cost, latency, stability, and resource elasticity of API calls can become bottlenecks. That's when inference service-ification is needed — optimizing at deeper layers like model routing, inference engines, caching, scheduling, and hardware adaptation.

Q11: Will cloud services in the AI era shift from selling resources to helping applications validate their unit economics?

Yao Xin: Yes, this is one of the most important changes. Currently, AI applications with clear commercial closed loops are mostly enterprise-facing. But this To B differs from the past — it often starts with individual users within enterprises. We call these people "professional consumers," or Prosumers. They have professional needs, technical capability, and willingness to pay. AI efficiency tools solve the problem of "saving time," so the business math is easier to calculate. Therefore, an AI inference platform's value proposition can't just be "we have cheap inference services." It should become "we help you lower per-task cost and validate unit economics." Traditional cloud solved "compute fast." The Agent era needs to solve "actually get the work done." Infrastructure no longer sells just tokens but cost-structure optimization and a complete Agent work environment.

Q12: What do customers actually buy from PPIO?

Yao Xin: To put it simply, mainly three things. First, model API. PPIO connects mainstream open-source models and multimodal capabilities in language, image, voice, and video through a unified API. Developers don't deploy models themselves; they pay by actual token consumption. The platform currently handles over one trillion daily token calls and has over 560,000 registered developers. Second, Agent sandbox. This is runtime infrastructure specifically designed for AI Agents, providing a secure, isolated cloud environment where Agents can execute code, access file systems, and call networks. Agent-generated code is uncontrollable and can't run directly in production. PPIO's Agent sandbox business grew over 100x in scale in the year since launch. Third, GPU services and enterprise private deployment. For large customers with compliance requirements or dedicated resource needs, we provide GPU container instances and enterprise-grade private deployment solutions. Ultimately, developers don't need "compute"; they need "a service that runs": models callable, Agents runnable, enough elasticity, and a reasonable price.

Q13: What practical results can PPIO deliver in inference cost, latency, and throughput?

Yao Xin: PPIO's inference optimization covers multiple layers, including operator optimization, inference engine optimization, storage system optimization, and model adaptation for cutting-edge hardware. Through these optimizations, we push compute resource utilization toward theoretical limits, continuously reducing inference cost. Using 2026 data as an example, input token prices for some well-known large models delivered through PPIO's model API service are over 40% below the average prices of major international cloud providers. When we evaluated a well-known large model, PPIO-optimized service, compared to directly deploying the same model with the SGLang baseline inference engine, achieved a 95% reduction in first-byte latency and a 300% improvement in total throughput efficiency.

Q14: What is PPIO's relationship with large model companies, application companies, and traditional cloud vendors?

Yao Xin: With large model companies, it's an upstream-downstream relationship. They build the engines; we handle model hosting, intelligent routing, and inference optimization. Last year the open-source model ecosystem was less rich, but this year Chinese open-source models have gradually reached consensus. From Zhipu and DeepSeek to Tongyi Qianwen, open-source model capabilities continue to strengthen, which has also energized independent cloud platforms. With application companies, it's a service relationship. We hope to support today's "Yao Xins" and accompany them to success. With traditional cloud vendors, complementarity exceeds competition. Big vendors also provide AI cloud services, but usually tie them to their own cloud ecosystem or models. PPIO, as an independent third-party platform, provides neutral model selection, cross-model routing, and inference optimization, not locked to any single cloud vendor or model company.

Q15: What will be the core moat of the AI infrastructure industry in the future?

Yao Xin: First, quality AI compute resource reserves. The platform needs enough high-performance GPU resources, with dynamic inference scheduling and off-peak resource sharing to meet high-concurrency and real-time performance requirements. Second, inference optimization capability. The platform needs to continuously improve model speed and reduce cost through operator-level optimization, inference engine design, memory management, and hardware adaptation. Third, API compatibility and platform generality. It needs standardized APIs, stable SDKs, and universal platform tools so developers can deploy across models and scenarios quickly. Fourth, end-to-end Agent infrastructure capabilities, including secure execution environments, state management, memory systems, and workflow orchestration. Simply reselling compute is a business but hard to build into a long-term venture. The real moat is whether you can provide complete inference service capabilities covering the full lifecycle of AI applications.

Q16: From entrepreneur to investor and back to the entrepreneurial frontline, what different judgment do you have about today's AI entrepreneurs?

Yao Xin: I must admit I may not be able to return to the state I was in at 24. Therefore, we choose to support today's "Yao Xins" and accompany them to success. The timing is right for building an AI inference platform today. A basic industry consensus is that the iteration speed of model capability is slowing relatively, and AI is entering the application era. Last year Silicon Valley focused on training; this year the focus has shifted to inference. Next year is expected to be the Year of the Agent, when masses of Agents begin taking over existing workflows. There remains a massive infrastructure gap between models and applications — that is PPIO's opportunity. I view the AI entrepreneurs on our platform with an open, appreciative, and respectful attitude, even those 10 or 20 years younger than me. Among them lie new opportunities and hope.

Q17: What do you think of the current "heat" and "bubble" in AI entrepreneurship?

Yao Xin: My judgment is simple: AI that doesn't talk about revenue implementation is a bubble. AI coding, AI Agents and enterprise automation, AI video and content generation, and AI education are all directions with relatively clear willingness to pay and efficiency gains. AI coding is especially important because code itself means automation and intelligence. In the past, programming was considered high-IQ work; when AI can generate code, it means it can not only answer questions but further control computers and complete tasks. What's easier to commercialize at this stage are scenarios where accuracy requirements aren't extreme but pain points are clear and willingness to pay exists — such as marketing, meetings, and enterprise back-office tools. These scenarios, first adopted by Prosumers, are the first to commercialize.

Q18: Over the next three years, what will be the biggest change in AI infrastructure?

Yao Xin: Per-token prices may drop one to two orders of magnitude, but token usage will grow four to five orders of magnitude. At the same time, the shift from Chat to Agent will amplify compute needs for the same user scale by two to three orders of magnitude. Generative AI is "short spikes" — one inference completes in seconds and resources are released quickly. Agentic AI is "long duration" — one task may run for minutes or even hours, continuously calling the model. This isn't a short-term peak but sustained, high-concurrency inference demand. In the future, inference will replace training as AI's primary resource consumption scenario, and AI inference platforms will shift from "selling tokens" to "selling Agent work environments." Masses of Agents will begin taking over existing workflows, and the AI capability obtainable for the same spend may improve tenfold every year. This trend will directly spawn a batch of free or nearly free AI applications. The truly important infrastructure of the AI era is not piling up more GPUs but making every token call smarter and cheaper. This is consistent with the essence of technology development: continuously improving the efficiency of resource utilization.

Originally published by Unique Research on Unique Research Substack on July 15, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research