
Original · Unique Research · 2026-08-13
Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening essay, five insight sections, closing reflection, and the full roundtable transcript. All named companies, products, and people are preserved. Internal metrics, customer cases, and billing figures are speaker claims, not independently verified findings.
AI Industry Observer
A 300-Person Company Running 900+ Agents
"Whether the model is smart enough is just the entry ticket. Trust, accounting, and business model are what determine whether an Agent can walk from the demo room into the production server room."
99.9% single-call success rate — sounds rock solid.
At this roundtable at the Unique Awards, PPIO co-founder and CEO Yao Xin (姚欣) did the math on the spot: a task executed repeatedly 20 times, each loop calling a large model once — the failure probability gets amplified dozens of times, and the final failure rate could reach 2%.
That 2% is unacceptable in many serious work scenarios. Transactions, services, customer service — none of them can accept it. Yao says this is the core reason limiting whether Agents can land from ideas today.
You see, the problem suddenly changes flavor. When people usually discuss Agents, they talk about whether the model is smart enough, whether context is long enough. But at the door of the enterprise production line, what blocks the way is a string of dry numbers: failure rate, cost, bills, billing units.
This roundtable's theme was what kind of foundation enterprise-grade agents need. The host was Huang Shufei from PraxisGrowth. The four on stage are all infrastructure practitioners: David from SandBase.AI, Pan Yong (潘勇) co-founder of Molink, Li Li (李粒) from TiDB's AI business, and Yao Xin co-founder and CEO of PPIO.
After listening, I had a strong feeling: the point technologies of the Agent foundation are already ready. What's not ready are three unsexy things — trust, accounting, and business model.
Being Connectable Is Only 60 Points: Agents Need ID Cards Too
The host posed a very relatable scenario: your Agent performs perfectly in the demo, connects to real business systems, and the next day pushes an already-offline product to a client based on yesterday's inventory data. The day after, it forgets the process approved that morning. The day after that, the security team discovers it looked at client information it shouldn't have, without leaving any operation record.
The model is very smart. Where's the problem?
Molink's Pan Yong answered directly: it's trust.
The demo pursues "it can run." But for enterprises, the question is whether they dare use it, and whether problems can be covered. He said: no matter how well something is made, without identity, without auditability, without traceability, without accountability — no matter how beautiful, it can't enter production.
"Note what he said next — I think it's one of the clearest judgments of the session: enterprises don't worry about Agents making mistakes, because employees make mistakes too. What enterprises really worry about is whether, after a mistake, they can block it, find it, trace the cause, and assign responsibility to a personified subject. In plain terms: an enterprise can accept an employee who makes mistakes; what it can't accept is a black box that made a mistake and left no trace of who."
There are already many protocols on the market. MCP solves Agents calling tools; A2A solves Agents talking to each other. Pan Yong's critique was blunt: these all solve "being connectable," and being connectable is just the passing grade for Agent communication — 60 points.
"We met and connected — why should I trust you, and why should you trust me?" Without this, everything afterward is empty.
So Molink's AUN protocol's core is issuing each Agent a globally unique, verifiable, distributed identity ID. Audit answers four questions: who am I, what can I do, what have I done, who is responsible for me.
One detail left a deep impression. Pan Yong said they position Agents as subjects, like people, not endpoints. You send a friend a WeChat message and they don't reply — you think that's normal; you send an Agent a task, and it also judges whether it can do it and whether it should be done by it.
"It's a network citizen — not the kind that only works when you call it and doesn't think when you don't."
"The first half, everyone competed on capability; the second half starts competing on trust. Pan Yong dropped a line: identity, audit, authorization — today they look like bonus points; next year or the year after, they'll become entry barriers."
The Fat-Container Debt, Agents Don't Want to Pay It Again
David's SandBase.AI does Agent runtime. He answered a question on many people's minds: enterprises already have containers, serverless, and cloud VMs — why do Agents need a dedicated runtime?
He said it's not that resources changed, but that the object using resources changed. Originally, the ones creating monitoring, reading logs, and following processes were humans; in the future, it's Agents. Humans have brains, they filter, they go through tickets; Agents also need a process, and key points also need approval.
What's really interesting is his retrospective of industry history. From bare-metal hosts to K8s, everyone experienced fat containers: stuffing a bunch of services and state into a container, then finding it didn't work, and starting to emphasize statelessness — state either persisted or went into a database.
The Agent era is reenacting this. Anthropic proposed Claude Managed Agents in April this year, with the core idea of separating sandbox, runtime, and session. The loop engine manages context and stateful storage; the sandbox that actually executes tasks should ideally be stateless — same as a K8s pod.
"Anyone who's done infrastructure knows that architectural debt is cheaper the earlier you pay it. The Agent circle is now paying the tuition that cloud native already paid, except this time someone already wrote the answer ahead. David also said something honest: he insists on open-sourcing the runtime because to-B enterprises first demand control. It's exactly the same psychology as when enterprises first adopted the cloud. Technology can be advanced, but trust must be privately deployed."
A Million Databases, and an Account That Must Balance Debits and Credits
TiDB's Li Li pulled the conversation back to earth: money.
"We four all do infrastructure. Infrastructure often answers one question: do your costs and revenue, your margins, actually support making this a business?" He said almost every client discussing a new project starts with: how much will this cost, can we bear it?
He shared a concrete case. Minds uses TiDB for most of its infrastructure; their core request was: how to provide over a million databases to Agents every day at low cost, with each Agent's storage completely isolated — wrong token means absolutely no access to others' data.
Million-scale, cheap, completely isolated — these three words together are the data layer's new requirement for the Agent era. Previously databases were designed for humans, and the database itself was a scarce resource; now Agents outnumber humans, and databases must be as cheap and hygienic as disposable chopsticks.
Even more striking is the following example. Inside TiDB itself, the manual service bills generated for clients are produced by Agents. Li Li says the key here is treating the Agent like a person.
Imagine a person generating manual bills — do they need to verify? Do they need three-way data reconciliation? As finance professionals, calculating only revenue without balancing income and expenditure is definitely unreliable; debits and credits must balance before you can say this period's billing is done.
Agents doing serious business are the same — they need multi-source cross-validation.
"Never think of Agents as invincible things you can just throw into a workflow and let run — absolutely not."
The host asked a tricky question: enterprises say they care most about security, but when actually purchasing, what's the key metric that really determines whether they pay?
"Li Li recalled: last year clients were asking, can you help me get things done. This year clients ask, how many people does using this save. The boss's brain is definitely doing the math. Saved three people — what are their salaries? How much is token consumption? If the math doesn't work, they keep looking. Very real, that's exactly how it is."
He also pointed a path for everyone doing AI: go overseas and earn USD. Same clients — Monica charges USD, Kimi charges RMB, monthly fees are similar, just the unit changed, and margins and cash flow are completely different worlds.
Traditional Subscription Is Dead: Agents Are Scalping the Cloud's Business Model
The most aggressive force in the room was Yao Xin.
The host asked him: Agent tasks are longer, stateful, and bursty — what in cloud infrastructure most needs restructuring? Compute, scheduling, network, storage, billing, or observability?
He didn't pick any of them.
There's a big improvement from CPU to GPU, but none of these are fatal. In the AI-native era, the most fatal change is that the business model and billing model need to be completely overturned.
He first stripped off SaaS's underpants. Earlier this year there was a saying that subscription is dead. The subscription model is inherently designed for humans; doing to-C products requires studying human nature — laziness and greed for small bargains. The essence of subscription is resource overbooking: only 10 units of resource, sell 20, bet that you won't use it all. Phone unlimited plans, broadband packages — all this logic.
But after January 2026, this model was punctured.
Anthropic previously offered a subscription coding plan; in March they said no more. Why? One person using the coding plan isn't scary — what's scary is OpenClaw using the coding plan, scalping too hard.
"This is the essential difference between humans and Agents. A human starting a VM rents it at least by the hour; what does an Agent running a Sandbox look like? Half an hour asleep, then the startup fires off a bunch of tasks, finishes in tens of seconds, sleeps for another 29 minutes, then does it again. What cloud vendor backends see isn't peaks and valleys — it's all spikes sharp enough to cut. And humans only use it during the workday; Agents use it 24/7. When the user object switches from human to Agent, four changes hit at once: fragmentation, high frequency, pulsing, and 24/7 — the overbooking model is directly scrapped."
The business model changes, which drives billing model changes, which drives sales model changes, which drives virtualization and resource reuse methods to all change. Yao says this is truly unprecedented. On PPIO's website, Sandbox billing is down to 0.00000001 RMB per CPU per second.
One-millionth of a yuan per second. You might ask how you collect this money? Honestly, I want to know too, but the direction is clear: when the customer becomes a machine, pricing must be precise to the machine's time granularity.
On payment, Yao's judgment is the sharpest: clients don't pay for technology, don't pay for resources — clients always pay for the final result.
He says this with 22 years of industry experience. Back when he did PPTV, he witnessed four transitions in advertising billing: earliest was CPT — how much for a front-page headline on Sina for a day; then the video industry pushed CPM — how much for a million impressions; then it was killed by search engine CPC — pay per click; finally games and e-commerce moved to CPA — take a cut of goods sold.
From per day, to per impression, to per click, to per transaction. The billing granularity keeps approaching the result.
"All current discussions about Agent pricing are transitional states. Eventually you pay based on how much money the Agent made the enterprise boss — pay for results. It will definitely reach this point. Pay-for-results means no one can collect money by telling stories; whether the foundation is good, the ledger will speak for you."
Human in the Loop vs. Human on the Loop: Token Consumption Differs by 1000x
In the final advice segment, Yao said the hardest part of deploying internal enterprise Agents isn't technology — it's the AI-native organization. PPIO itself pushed AI for AI internally for a year: 300+ people running nearly 900 Agents.
From running this, he found there are two types of scenarios. One is human in the loop — the human is inside the loop, like AI coding, where the most common experience is repeatedly clicking yes, confirming, deciding. This can run, but the ceiling is obvious.
The other is human on the loop — the human is watching from above the loop. Even when you're asleep, a pile of Agents are still working, serving, iterating.
How to tell if a team is truly AI-native or just pretending? Yao gave a particularly hard metric: look at the token bill.
Business and teams that are on-the-loop have token consumption two to three orders of magnitude higher than other teams. It's not a 1:10 problem — it's 1:100 to 1:1000.
So his advice is specific: don't do a company-wide overhaul. First find the single-point business in the company that can truly self-loop, and pour resources into it. Catch one and it's worth it.
The host pressed: how did you discover the pain point when doing AI-native transformation inside the company?
Yao's answer made everyone laugh. "The first step is to get close to Gen Z, avoiding the participation of all you old-timers on stage. I first learned about tools like Marscode and PaperClip from the youngest, laziest, least experienced programmers. I'd talk to them, find they were doing this, learn from them, and promote it."
A founder who built PPTV 22 years ago and now runs a distributed compute company — his signal light for organizational change is the laziest young people in the company. Makes sense, when you think about it.
"Laziness is the original driving force of human progress, and Agents are essentially the serious solution prepared for people who want to be lazy."
More Conversation Details
Speakers
SandBase.AI Founder & CEO Li Yangbing David (李样兵)
Molink Partner Pan Yong (潘勇)
PingCAP TiDB APAC AI Business & Product Leader Li Li Lux Li (李粒)
PPIO Co-founder & CEO Yao Xin (姚欣)
Host
PraxisGrowth Co-founder Huang Shufei Sophie (黄姝菲)
From Demo to Production: What Is the Real Dividing Line?
Huang Shufei: I'm Sophie from PraxisGrowth. We run an AI-native GTM accelerator for going overseas, helping Chinese AI entrepreneurs going overseas be seen by the world, and helping overseas investors and entrepreneurs enter China. We're a bridge and hub. It's an honor to discuss today's topic with four founders: what kind of foundation does an enterprise-grade agent need?
Huang Shufei: Let's start with David — please briefly introduce your business.
David: Hello everyone, I'm David. Our product is SandBase.AI — AI Agent infrastructure, hoping to let builders and FDE engineering teams quickly deliver Agents. That's our vision.
Pan Yong: Hello everyone, I'm Pan Yong, co-founder of Molink. Molink is an Agent infrastructure service provider. We released an Agent communication protocol called AUN, defining communication between Agents. Based on AUN, we also built a product that gives Agents identity, so they can be trusted and invoked across organizations. This is our latest product, launching soon.
Li Li: Hello everyone, I'm Li Li from TiDB, responsible for TiDB's AI business and product. TiDB mainly provides data layer and infrastructure layer for Agents and base model vendors — various state layers, data layers, and data interoperability, empowering all Agent vendors. We're also increasingly working with small enterprises — feel free to reach out if interested.
Yao Xin: Hello everyone, I'm Yao Xin, co-founder and CEO of PPIO. PPIO's core services are two: one is AI cloud — what everyone calls the token factory today, mass-producing tokens; the second is Agentic Cloud, providing services from Agent Sandbox to Agent application platforms. I'm a serial entrepreneur — 22 years ago I built PPTV, the first video software, which became a distributed streaming media distribution system; today I've built a distributed compute system, integrating CPUs and GPUs into Sandboxes and tokens for developers.
Huang Shufei: When Agents actually start reading data, calling tools, modifying systems, and executing tasks, how do you make it run stable, manageable, traceable, and accountable? In this next segment, let's quickly establish a shared definition: from demo to production, what is the real dividing line? If you could only choose one production red line — data consistency, model availability and security, execution permission audit and task success rate cost — which would you choose in an enterprise, and why?
David: Sophie asked this well. The road from idea to production is still quite long. The stability, audit, and permission issues mentioned are all important. But I think the more important point now is how to string these together. There are two groups: one is builders, whose core need is to make money fast. They have skills and scenario understanding, wanting to quickly turn skills into revenue-generating APIs or SaaS — like building an MCP service that Cursor and OpenClaw can quickly integrate. That's the builder's need. For builders, security and audit matter, but they're more focused on how to quickly run the demo, get the flow working, find the first wave of customers, and then improve the infrastructure further.
For enterprises, security is more important. Internal security audit matters, but the more important point now is that enterprises haven't actually started deploying Agents to production yet — many things aren't ready for this step. People may have started by running an SDK with LangChain, but running it is just the first step. How to keep this service running continuously, continuously providing production-grade service — that's very important. Everyone is exploring how this path should go and whether there are best practices. Looking at the entire Agent runtime system, from Anthropic proposing Claude Managed Agents on April 8, 2026, to OpenAI releasing the Agents SDK on April 9, including Cloudflare following up on these protocols and implementations — but actually, up to now, I don't think there's a relatively good practice or process; everyone is gradually exploring, but there should soon be a relatively unified view on what environment issues exist from demo to production, how to solve them, and how to start.
For enterprises, beyond security audit, the most important thing is how to start. Having an idea, you can't just deploy something like OpenClaw — with little security control — directly to production; the risk is high. We need something that lets enterprises deploy to production — like a best practice or an initially agreed-upon standard protocol — so people can practice and quickly run it. Then solve the security and audit issues inside.
Pan Yong: In my view, the demo pursues being runnable and achievable — that's just the first step. For enterprises, the problem is whether they dare use it, and whether problems can be covered. So the difference between demo and production mainly comes down to trust — dare they use it. No matter how well something is made, if it doesn't hold up to scrutiny — no identity, no auditability, no review, no traceability, no accountability — no matter how beautiful, it can't enter production. For enterprises, controllability, trustworthiness, and traceability are very important. Enterprises may not worry about using Agents making mistakes, because employees also make mistakes. What they really worry about is whether after a mistake, it can be blocked, found, traced to the cause, and responsibility traced to a personified subject. That's the logic that truly pushes products into the enterprise production line.
If I could only keep one red line, I'd choose permission audit. Because data consistency, model availability, success rate, and cost are about "doing things right"; permission audit is about "when things go wrong, you can check, block, and hold accountable." The former determines whether the Agent is good to use; permission audit determines whether enterprises dare to use it. A system you don't dare hand over can't enter production no matter how right it is. These points — the others ensure things can be done, but audit and review ensure enterprises dare use it.
Li Li: What Mr. Pan just said is very right — trust is very important in the path from demo to production and must be solved. Another thing: Agent systems, like our normal internet systems, the simplest is whether your costs, revenue, and margins support making this a direct business. We're four infrastructure people here; infrastructure often answers this question — what kind of infrastructure is most suitable for each person's Agent system, what data storage method is best, how to store it most cost-effectively with highest reliability, and how to use it for future Agents without leaving pitfalls. This is definitely the core.
Many clients we serve, when exploring new projects, ask: how much will this cost, can it bear the Agent business? Every Agent launch requires calculating whether the money is enough. So cost is very important, and we're also working to provide the cheapest Agent infrastructure.
Yao Xin: I'll pick a different one — I choose success rate. Why? Don't forget the underlying principle of Agents and large models. Large models differ from all previous app development in that the underlying layer is a probabilistic system with hallucinations. Even the most advanced large models today have some failure probability, but at the Agent layer, that failure probability gets infinitely amplified. For example, in loop engineering, when a task executes repeatedly, if the underlying model success rate is 99.9%, one-in-a-thousand failure is already good, but after 20 loops the failure probability increases dozens of times, and the final failure rate could reach 2%. That 2% is unacceptable in many serious work scenarios — many transactions, services, customer service can't accept it. So I think this is the core reason limiting whether Agents can land from ideas today.
When Agents Go Wrong in Production, Where Exactly Is the Problem?
Huang Shufei: What is the core contradiction of enterprise-grade AI? Is it whether you have a large model? Let me pose a scenario — your Agent performs perfectly in the demo, answering client questions and searching the knowledge base. After connecting to the real business system, the next day it pushes an already-offline product based on yesterday's inventory data; the day after it forgets the process approved that morning; the day after that the security team discovers it looked at client information it shouldn't have, without leaving any operation record. The model is very smart, but where is the problem?
Li Li: Our clients have actually encountered this exact problem — not alarmist. What was just mentioned actually consists of several problems; I'll briefly say two angles.
First, every Agent running needs a completely isolated environment. Mr. Yao has a very good sandbox; each Agent's independent runtime environment must be completely isolated — not just compute-time isolation, but storage-time isolation as well. This is also what we're working on. For example, everyone knows Minds — most of their infrastructure uses our database. One of their requests is how to provide over a million databases to Agents daily at low cost — millions per day. In this case, the infrastructure problem they raised is: they want each Agent's key information and storage to be a completely independent space — as long as you access with a token and the token is correct, what you access absolutely won't mix with others'. This is very important. Whatever we do — file systems, databases — we can provide million- or ten-million-level independent, cheap storage specifically built for Agent products.
This is the first layer — reading others' data can be completely prevented at this layer. Of course, if the LLM swaps one ID for another, there's no good solution for now — let's hear other experts' views.
Second, the transfer-wrong-money example mentioned earlier isn't actually an infrastructure problem. Our own internal business also has critical business decisions — the manual service bills for clients are generated by Agents. Clients buy human services; human service statistics come from all over; many clients need this bill, generated by the Agent and dispatched by the Agent to the quality system. In this process, you must treat the Agent like a person. The Agent isn't omnipotent — never think of Agents as invincible things you can just throw into a workflow and let run. Imagine if a person generated a manual bill — would they verify? Would they do three-way data checks? As finance, calculating only receivables without balancing income and expenditure is definitely unreliable; expenditure and revenue must be balanced, debits and credits must balance before you can say this period's billing is done. Agents doing serious things like bills are the same — they need multiple data points to validate. This is a problem the business must solve, and we're solving it ourselves.
What Core Gap Does the AUN Protocol Hope to Fill?
Huang Shufei: Second perspective, from Mr. Pan at Molink. Molink recently released a new ACP protocol; the market already has MCP, A2A, and various Agent collaboration protocols. From the enterprise production environment perspective, a universal Agent protocol must solve not just connectivity but identity authentication, capability discovery, authorization, and security audit. I'm very curious — what core gap does the ACP protocol hope to fill?
Pan Yong: Here's how it is. There are many protocols on the market now, including MCP and A2A. MCP solves the problem of Agents calling tools; A2A solves Agents talking to Agents. Both solve "being connectable." But being connectable is just the basic passing grade, 60 points. The real question is: we met and connected — why should I trust you, and why should you trust me? This is especially important. Without this, everything afterward is empty and foundationless.
From this angle, what AUN tries to solve is giving each Agent a globally unique, verifiable, distributed identity ID — this ID is unique and verifiable. It's like a company with many departments; every department employee's card must be unified — only then can cross-department collaboration authorize employees, and problems can be traced. I've been emphasizing identity, authorization, and audit today — these run through this line. We want Agents to communicate, and the purpose is to hope for chemical reactions and deeper collaboration; the premise of collaboration must have a foundation of trust. With a trust foundation comes an ID, a unique identity — traceable, verifiable, following it all the way, stringing the whole process together. Only then does future Agent ecosystem collaboration become reality. This is the core point AUN tries to solve — the identity problem.
Huang Shufei: Understood — Agents also need identity and identity recognition.
Pan Yong: Right, let me add one point. We see Agents as subjects, not endpoints or services. Traditional services — I ask, you must answer; if you don't answer, I think you timed out or said it wrong. Agents are different; we position them as subjects, like people. You send a friend a WeChat and they don't reply — you don't think that's abnormal, do you? No. Agents are the same — they have role boundaries; after you send something, it judges whether it can be done and whether it should be done by it, with autonomy. It's a network citizen — not the kind that only works when you call it and doesn't think when you don't.
Why Does an Agent Need a Dedicated Runtime?
Huang Shufei: Next question for David. SandBase.AI has always emphasized Agent-native runtime and sandbox, but enterprises already have containers, serverless, and cloud VMs — why do Agents need a dedicated runtime? In long-task state preservation, code execution, browser operation and tool permissions, failure retry, and log tracing — what exactly is missing from traditional cloud infrastructure? Why is such a new concept needed?
David: Good question. Many people equate Agent runtime with sandbox — earlier ones like North America's E2B came early, and many people's first contact with Agent runtime was through sandbox. Many people now say sandbox is Agent runtime, which is a bit narrow. For runtime, beyond sandbox, there are many things to solve. Li Li just mentioned audit, permissions, security — these are all problems. Originally people's understanding was scattered, and after these things existed, individual point capabilities are ready. Just mentioned container hosts — point capabilities are ready.
Agent runtime isn't replacing the original Docker or containers; it's better stringing together resources and processes. What runs underneath Agent runtime is process isolation, Docker, user hosts — the resources are the same, they don't change. Including user business processes — originally using monitoring, log tracing, these don't change either. What changes is that originally these things were used by humans — creating monitoring, creating logs, reading these things; in the future, it's actually through Agents. Humans raise requests and demands; Agents do the work. How does the Agent better use infrastructure — it's not that the Agent replaces infrastructure, but that the Agent can better use infrastructure, and the interaction method will change. After all, it's not human — humans have brains, need to filter, process things, go through tickets and processes. For Agents, do they also need a process, with key points requiring approval? I think they do.
How to let Agents better push the process forward and run it — that's the core. In the process, infrastructure doesn't change; Agents go string those things. More importantly, we need a good runtime protocol and standard to string everything together.
Why do I keep mentioning Agent runtime? Originally when people mentioned Agent runtime, it was scattered — either sandbox or link tracing logs. But Anthropic proposed a good concept — Cloud Managed Agents — the core is separating sandbox, runtime, and session. Why separate? Some early solutions like OpenClaw are relatively fat containers at many points. From bare-metal hosts to K8s, everyone experienced the fat container concept — putting many services and states into containers. In K8s, people emphasize statelessness; state either persists or goes into TiDB or a database. The Agent era is the same — separate state from the actual loop engine. The loop engine is responsible for context or some stateful storage; the actual execution sandbox is hoped to be stateless. This aligns with the original infrastructure — like a pod has no state, Agents also hope to have no state. State is stripped from the original workspace and can be used in memory and storage. The underlying sandbox — like what PPIO is good at — sandbox tuning capability doesn't change.
I think if we clarify the boundaries, everyone has a set of aligned cognition, then we can talk about how to do delivery, how to do enterprise FDE delivery, and how to let enterprises use it with more trust. First unify boundaries and cognition, then talk about delivery. I started doing Agent runtime from open source because many things are cognition problems — people need to recognize that this thing is controllable. If the underlying storage uses MySQL or OSS and is theirs, they feel it's controllable. If you provide a set of things for them to use, they're probably not that reassured — no matter how well you say it, they think it's not safe. It's like when users first moved to the cloud — in the Agent era, to-B enterprises definitely want to be controllable first. You can tell them it uses sandbox, host, logs — at least it's their own thing, responsible for getting things running first — that's the first step: let enterprise users start using it.
Everyone first has a common understanding of how to use Agent runtime. With this understanding, many things take time to solve — it's not that you can propose something now and immediately deliver or implement; there's definitely a road in between. In the process, we hope to take on the role of evangelist, telling everyone what best practices or initially unified standards are, then talking about how to use it better. But in the process, each ecosystem's original capabilities remain original capabilities — we're just telling everyone they can use it this way, string it this way.
SandBase's philosophy is more about doing evangelism, doing promotion, everyone building the ecosystem together, each doing what they're good at. We might be good at to-B and enterprise delivery understanding. Everyone goes back to their positions, unifying cognition and understanding of Agent runtime in the process.
How Is Agent Cloud Different from Traditional Cloud?
Huang Shufei: Next question for Mr. Yao. How is Agent Cloud different from traditional cloud? From PPTV's large-scale streaming distribution, to PPIO's distributed cloud, to now Agent Cloud — you've been through several generations of compute workloads. Compared to traditional web and mobile internet applications, Agent tasks are longer, stateful, bursty, and have uncertain execution paths. To support such workloads, what in cloud infrastructure most needs restructuring — compute, scheduling, network, storage, billing, or observability?
Yao Xin: First, I estimate everyone here is a programmer or technical person. Technically, there's a big improvement and change from CPU to GPU, but I don't think these are fatal. In the AI-native era, the most fatal change is that the business model and billing model need to be completely overturned.
Earlier this year, everyone knows SaaS companies faced a challenge called "subscription is dead." The subscription model is inherently designed for human characteristics; back then doing to-C products required studying human nature — one aspect is laziness, another is greed for small bargains. The subscription model is essentially resource overbooking — betting that you won't use all the resources allocated, so the cloud vendor can make extra money. Only 10 units of resource, sell 20 — phone unlimited plans, internet plans, all like this. But this model was challenged after January 2026.
Anthropic previously offered a subscription coding plan, then in March said no more. Why? One person using the coding plan isn't scary — what's scary is OpenClaw using the coding plan, scalping too hard.
Today's observation: when cloud users switched from humans to Agents, dramatic changes happened. Previously humans starting a VM rented at least by the hour; today how does OpenClaw run a Sandbox? Everyone knows there's a half-hour startup mechanism — asleep for half an hour, then the startup fires off a bunch of tasks, finishes in at most 30 seconds to a minute, then goes back to sleep for 29 minutes before doing it again. What the backend sees isn't peaks and valleys — it's all very sharp pulse-type applications.
When the user object of cloud and applications switches from human to Agent, usage becomes more fragmented, high-frequency, pulsed, and 24/7 continuous. Humans only use it during the workday; Agents use it 24 hours. The biggest impact of these characteristics is on the business model definition — the previous cloud overbooking model and SaaS subscription model are dead; they must be iterated and innovated with new business models. This is what troubles us most now.
Go to PPIO's website and look at Sandbox subscription — the subscription unit can be one CPU, one second, charged at 0.00000001 RMB. Of course there's the problem of how to pay that one-millionth of a yuan — in practice you'd never use just that little; you have to top up, etc. But the whole series of challenges is that Agent usage has completely changed.
Because the business model changes, it drives billing model changes, which drives sales model changes, which drives technical underlying model, cloud reuse model, virtualization method — all must change. This is what's currently involved — truly unprecedented, requiring re-research.
What Is the Key Metric That Ultimately Decides the Purchase?
Huang Shufei: Understood. A question just emerged, following up on Mr. Yao's business model answer. For example, enterprise-grade Agents say they care most about security, but in actual delivery, what do you find is the key metric that ultimately decides the purchase? What kind of Agents and services are clients more willing to pay for? In your back-and-forth with clients, can you say the honest truth?
Yao Xin: Let me be honest first: clients don't pay for technology, don't pay for resources — clients always pay for their own final results. I have this experience — I went through business model transitions. Earliest selling ads was CPT — how much for a Sina front-page headline for a day; then the video industry pushed CPM — how much for a million impressions; then it was killed by search engine CPC — how much for a million clicks; finally games and e-commerce went to CPA — today how much goods sold, what cut.
All current discussions are transitional states; eventually you pay based on how much money the Agent ultimately makes the enterprise boss — pay for results. It will definitely reach this point.
Li Li: Let me recall — very apt. Last year when talking with clients, they were still asking, can you help me get things done. This year they ask, how many people does using this save. The brain is definitely doing math — saved three people, what are their salaries? How much is token consumption? If the math doesn't work, they keep looking. Very real, that's exactly how it is.
David: Right, we serve enterprises both domestically and internationally. Clients do pay for results and business value. For the same product, where clients use it in high-business-value places, they're very tolerant of price; where used in low-business-value places, or places where traditional historical pricing was very low, tolerance is very low.
This is also why TiDB has mostly done overseas work, and we strongly encourage all AI companies to go out and see. Comparing our clients Monica and Kimi's monthly fees — one charges USD, one charges RMB, similar monthly fee just different units — you can immediately see the difference in margin and profit, and what the enterprise cash flow looks like. If you have the chance to go out, go out — try to earn USD, it's more comfortable. For infrastructure vendors, there's also more opportunity to provide higher-value services.
Pan Yong: From GPT's emergence in 2023 to now, enterprises' AI payment willingness has gone through three clear stages:
2023: Cold start period with no payers. Back then AI was hard to make money — except for course sellers, service providers basically had no revenue. The root cause was AI output quality was poor; bosses couldn't see value, naturally no payment willingness — "if the results aren't good, why should I pay?"
2025: Trial period paying for capability. The market started seeing AI actually solving problems in certain scenarios, and payment willingness broke through from 0 to 1. Enterprises were willing to pay for relatively clear single-point capabilities — API calls, basic service units, quantifiable things. The logic was "this capability is OK, I can pay for it."
2026 to present: Explosion period paying for results. After the Spring Festival, OpenClaw blew up and the real turning point came: people started willing to pay for results. Many enterprises actually saw cost reduction and efficiency improvement — labor costs significantly decreased, human efficiency ratios significantly improved. Especially traditional enterprises, when they saw peers bring costs down and efficiency up through AI, they immediately got restless and主动 came to talk, willing to pay.
This change is significant. In 2023 and 2024 when doing this, nobody wanted to talk to you; now at least bosses are willing to listen and try.
How to Make Agents Truly Enter the Business Closed Loop?
Huang Shufei: Finally, I'm curious to ask everyone: if you give one piece of advice to everyone at the Unique Awards — including enterprise decision-makers, AI entrepreneurs, investors — about how to make Agents truly enter the business closed loop rather than stay at the demo stage, what advice would you give? Each shares based on your business and actual scenarios. Starting with David.
David: For Agents to truly run, the underlying infrastructure they depend on is quite extensive; it needs many people and many vendors building together — it's not something one person can do. The infrastructure field still has quite a lot of opportunities; I hope everyone looks more at infrastructure opportunities, including in verticals or relatively doable areas — maybe one point, maybe one API — everyone together making the infrastructure ecosystem bigger.
Pan Yong: In the first half, everyone competed on capability; in the second half, trust will start being competed on. Today identity, audit, and authorization look like bonus points, but roughly next year or the year after, trustworthiness, controllability, traceability, and identity will become entry barriers and required items. Previously people didn't realize this part, but starting now you must clearly realize: you must know who the Agent is, what it can do, what it has done, and who is responsible — this is very important.
Li Li: Exactly echo what I said earlier — if you're still doing Agent business, you must go overseas and earn USD. The overseas market is really wide. If you can earn USD, really go earn USD and come back to pay taxes. If interested in going overseas to earn USD, contact me — I can give credits, introduce business, we have offices worldwide, can even do business promotion — several partners have done this. You can also use Mr. Yao's Sandbox.
Yao Xin: I have a different view. For technology and products, you can find us here, but for internal enterprise Agent deployment, the hardest part isn't technology — it's the AI-native organization. Our own enterprise is also pushing AI for AI internally, for about a year. From all-staff AI coding and all-staff programming, to today's large-scale internal Agent systems. 300+ people have nearly 900 Agents running.
We clearly found two types of scenarios: one is human in the loop, one is human on the loop. Most human-machine symbiosis scenarios can run under human-in-the-loop — like AI coding, the most common experience is clicking yes, confirming, deciding. But truly running efficient Agents — a native organization must be on the loop — even when asleep, a pile of Agents are still working, serving, iterating.
We provide token services ourselves and can clearly see that on-the-loop businesses and teams consume two to three orders of magnitude more tokens per month than other teams. It's not a 1:10 problem — it's 1:100 to 1:1000. Each enterprise should first dig out which business can truly reach self-looping loop engineering, and put effort into this single point — no need for full-process transformation; catching the core single point is more important.
Huang Shufei: When you were doing AI-native transformation inside the company, how did you discover that pain point? Can you share with everyone?
Yao Xin: First step is getting close to Gen Z, avoiding the participation of all you old-timers on stage. Getting close to young people — I first learned about internal systems like Marscode and PaperClip from the youngest, laziest, least experienced programmers. I'd talk to them, find they were doing this, learn from them and promote it.