
Original reporting by Unique Research · September 6, 2026, 20:00 · Shanghai
The real AI gap between China and the United States may be about payment, not models
.
AI Industry Watch
What the media calls the “global rankings” should really be called the “global rankings on the OpenRouter platform.”
You have probably seen some version of this line in the news: another Chinese open-source model has reached a new position in the global token-consumption rankings.
It sounds encouraging.
But in a livestream conversation with Yang Pan a few days ago, he challenged that framing on the spot: OpenRouter’s token consumption accounts for only 1% to 2% of actual worldwide token consumption, he said.
What does that mean? What the media calls the “global rankings” should really be called the “global rankings on the OpenRouter platform.”
The bulk of token consumption sits in contracts that enterprises sign directly with cloud providers and model companies—in places you cannot see.
Yang said he had checked that 1% to 2% calculation himself and cross-checked several serious data sources. In his account, that is broadly the industry’s consensus estimate.
Yang is a former co-founder of SiliconFlow and is now building an AI venture called Bilink. He has worked on the front lines of computing infrastructure, models, inference and applications; he brings a perspective that spans the industry chain.
Our conversation ran for two hours, covering everything from the first half of 2026 to AGI. There was a great deal to take in.
I have pulled together some of his sharpest judgments. They are worth ten minutes of your time.
Start with an example you can feel.
Yang said that asking an agent a question today, compared with asking that same question when DeepSeek first came out in early 2025, might consume ten times, a hundred times or even 1,000 times as many tokens behind the scenes.
Over the past year, the biggest undercurrent in the AI industry has been exponential growth in token consumption.
Users feel that “the money is burning too fast.” Providers see an entire market exploding.
In Yang’s account, that demand drove the turning points in GPUs, computing capacity and the global memory market from late last year into early this year.
Half joking, he said that if you had understood the trend in September or October last year and started paying attention to companies such as Samsung and SK Hynix ahead of time, the outcome might have been very different.
But during the first half of this year, another attitude was fashionable: Token Maxxing. The more tokens you burned, the better you seemed to be at using AI.
Yang’s verdict was blunt: this was a classic “peak of ignorance.”
People have now started sliding down from that peak into the valley of despair.
What started the slide? People gradually realized that consuming more does not mean producing anything valuable.
Some people used AI to mass-produce things with no commercial value. Online, this is called “AI slop.”
Around Lunar New Year, a number of large technology companies invited Yang to speak about “how AI can improve efficiency.” Six months later, the same companies came back, but the question had changed: “How do we actually unlock efficiency and economic returns?”
From showing off consumption to questioning returns: that is the most authentic emotional arc of the first half of the year.
Before discussing areas of consensus, we need to address a widely repeated claim: have models hit a capability wall?
Yang’s answer was unequivocal: not remotely. There is still substantial room for improvement.
He offered a perspective I had not paid much attention to before.
Last October, Claude Code introduced Skills. Most people understood them as “a more useful plug-in.”
But from the perspective of model evolution, Yang regards this as a remarkable invention.
Why? Once Skills appeared, large numbers of human experts began contributing previously private expertise in the form of a Skill.
Once a cloud model has used a Skill, some mechanism may give the model a chance to internalize that capability, he argued.
According to Yang, Anthropic published a blog post in the first half of this year dividing Skills into two categories:
One is “capability uplift.” It is as though humans have written an exclusive manual for the model: this version still needs to follow the manual, while the next version may have learned to do it on its own.
The other is “encoded preference”: working habits and preferences, such as always following steps one through five or using a particular template. These are not about capability.
The first category is the genuinely valuable one.
“Humanity is distilling itself, feeding its most closely guarded craft to models, one spoonful at a time.”
In Yang’s view, much of the improvement in models during the first half of this year has come from that dividend.
More recent explorations such as DeepSeek Harness, which combine models and tool frameworks for reinforcement learning, could bring another leap, he said.
At this point, a clear chain of absorption emerges.
First we wrote prompts, and models absorbed the prompts. Then people built workflows, and those were absorbed too. Now the Skills installed in agents are becoming built-in capabilities of the next generation of models.
Yang has a provocative prediction: models will eventually capture 80% of the economic gains in the global software and internet markets.
What, then, is the remaining 20%?
The things that cannot be trained into a model.
Real-time data, for example. A model can use probabilities to estimate what the weather in Hangzhou has usually been like on August 31. But to answer “Is it raining outside right now?” it must ask someone who has real-time data—and that person can charge the model.
Another example is the interface to the physical world. A model can do everything for you in the digital world, but the final hundred meters of a food delivery, or a machine in a factory, still depends on capabilities packaged in the physical world.
“The digital and physical worlds are separate. Anyone looking for an AI opportunity should write that sentence on the wall.”
Real-time data, the physical world and capability interfaces: in Yang’s framing, these are three things models cannot simply swallow, and they leave room for newcomers.
The first area of consensus concerns models. The second concerns people: agents have moved beyond programmers to serve all knowledge workers.
Consider a few numbers. By a broad estimate, China has around 10 million programmers. But it has roughly 100 million knowledge workers who sit at computers.
A tenfold difference.
Last year, people mainly used agents to write code. Claude Code, Cursor and Codex were programmers’ toys.
During the first half of this year, the baton passed to office and knowledge work. Claude Cowork led the way, Codex followed, and Chinese products including WorkBuddy, TRAE and Doubao Work rushed in. Large companies began committing computing and marketing resources at scale.
Yang has a cutting test: once something becomes a consensus, it enters the big companies’ comfort zone. When you start seeing it in elevator advertisements, you can be certain it has become consensus.
Looking back, there were several key moments in that expansion beyond the original audience:
DeepSeek’s emergence around Lunar New Year last year showed many people in central and other state-owned enterprises, and in non-IT industries, for the first time that “AI can help me with office work.”
Manus, in March last year, created the first global consensus about “what an agent looks like.” Before Manus, there was no shared understanding worldwide; afterward, people would say, “This product looks like Manus, so it is an agent.”
Then, early this year, OpenClaw—the “little lobster”—put autonomous agents onto millions of computers and gave the public a lesson in the process.
What drives each expansion? Getting people outside the original circle to actually use the technology. Enthusiasm confined to the technical community is not enough.
And why can agents suddenly get work done this year? There are two concrete indicators.
One is long-horizon tasks. People at Zhipu told Yang early on that their central measure of model progress this year was “how long it can keep working continuously after a single delegation.”
On the surface, this is duration. Underneath, it is the complexity of the work that can be completed in one assignment.
Give someone the same exam with a limit of 40 minutes, four hours or four days, and the work they submit will be completely different.
When a model is not intelligent enough, it cannot keep working longer even if you ask it to: it thinks it has finished and hands in the paper early.
The other indicator is hands and feet. Previously, models had a brain and could write programs, but they were weak at operating computers and browsers.
When a task reached “open the webpage, fill in the information and click Submit,” it had to stop and wait for a person.
Once Computer Use and Browser Use broke through that barrier, task duration increased substantially, because models were no longer constrained by “human speed.”
Yang offered an example from his own experience.
Last year, when he wanted to launch a website from scratch, he still had to do a long list of things himself after the code was written: register a domain, enter credit-card details, configure a server and connect Google Analytics.
This year, he said, he only has to say, “I want to build a website; help me check which domains are available.” The agent handles the remaining registration, deployment and submission to search engines itself.
AI’s real leverage lies in the time it saves you.
The internet reduced distances in space: you can attend a meeting without flying to the United States. AI compresses the time dimension.
If one AI can do a job and you have enough tokens, you can hire 1,000 AIs to do it simultaneously.
People cannot click a mouse while they sleep. AI can.
I found this the most counterintuitive part of the entire conversation, and it deserves to be unpacked.
In the past, making a product meant treating “create something, distribute it and earn the money back” as a single undertaking.
Creation required real money, hiring people and spending time, so you would only build something you believed had commercial value.
Today, AI has pushed the cost of creation close to zero, pulling those three stages apart.
“You have only completed the first step. The costs of distributing it and getting other people to pay have not fallen to zero.”
Yang calls the current state “compensation”: you once wanted to paint but lacked artistic training, and AI paints for you; you could not program, and AI writes the code for you.
Society is enjoying this compensation for things people once wanted but could not obtain. Completing them feels good and produces plenty of dopamine.
But that is precisely where the problem lies.
Nobody stops you from creating; you can do it at home. Market attention, however, is finite, and marketing costs real money.
That is why the world is now full of things that have been “made but have no value.” Creation certainly matters, but nobody is taking the next two steps for you.
It also explains another phenomenon: the “one-person companies” that have gained traction are concentrated almost entirely in areas such as independent media and AI video.
Those are areas where one person can relatively easily take the work from beginning to end and complete all three stages.
A deeper change concerns the software market itself.
Previously, when you had a need, you had to search for recommendations, download three to five programs and try each one. Often, you bought an enormous piece of software to solve a tiny problem. The fit never reached 100%, and the price was full of premiums.
What happens next? Everyone can use an agent to “hand-make” a piece of software for themselves whenever they need it, then discard it once the problem is solved.
The large-scale market for general-purpose software may shrink as personalized, disposable software grows: one gains as the other recedes.
Yang gave his own example. His computer had recently been running hot. Previously, he would have searched online for tools to identify the troublesome process.
Now he simply tells Codex, “My computer is overheating. Find out why and fix it.” Even the act of “finding a piece of software” disappears.
A related point concerns software interfaces.
Last year, Yang was an advocate of “generative software,” believing interfaces should be generated on demand. This year, he revised that view: the core issue is fulfilling intent, not the interface.
Voice, text or even an agent proactively recognizing what you need from context—use whatever is efficient. The interface is a secondary question.
This was the most uncomfortable judgment I heard. Anyone in management should listen.
If every employee uses AI, does that mean the company is good at AI?
“One person using AI might become ten times as efficient. With more than ten people, the gain might fall to just 20% to 30%.”
Why does the dividend become thinner as more people are involved?
Because the essence of an organization is collaboration, and the essence of collaboration is moving information from one brain to another.
That transfer uses speech, writing and documents. Every transfer introduces loss, misunderstanding and additional time.
Thinking something through in your own head may be ten to a hundred times faster than bringing someone else into a meeting to get them up to speed.
However powerful AI becomes, it does not close that “between-brains” gap.
So what can you do?
Traditional organizations divide work by job: product people understand product, developers understand development, and operations people understand operations. Limited human capabilities forced us into that form of collaboration.
But if one person can handle product, development and marketing with AI’s help, the entire way people are deployed can change.
The industry is already exploring several approaches.
One is a front-office/back-office model: the back office builds the thing as a self-contained unit, while the front office handles operations, marketing and results. The fewer points requiring repeated exchanges of information between them, the better.
Another is a role structure Yang said Anthropic had shared internally. Instead of traditional jobs, it divides work into roles:
The Prototyper creates prototypes quickly.
The Builder does the actual construction.
The Sweeper trims AI-generated slop into something reliable.
The Grower handles growth.
The Maintainer handles long-term maintenance.
Notice that each role owns a complete piece of work that can be delivered independently. The aim, in each case, is to minimize crossings between brains.
Behind this is the relationship between productive capabilities and the way production is organized.
Software development began with the waterfall model. Its most extreme form is construction: every step must be fully planned from day one because you cannot restart midway without enormous cost.
Agile development came later, as productivity improved two- or threefold and the cost of trial and error fell. You could make a prototype and change it if it did not work.
Today, AI has pushed that cost much lower again, close to zero. That is what makes a role such as Prototyper possible.
So stop asking, “What is the standard answer for an AI-native company?” There is no standard answer.
One axis is the nature of your business and its safety requirements. The other is your product’s stage: when searching for product–market fit, assign more Prototypers and Growers; once the system is stable, assign more Sweepers and Maintainers.
Every company must find its own path.
One exception is worth mentioning: work agents often produce more obvious efficiency gains for ordinary knowledge workers than for programmers.
Programmers can become more tired after adopting AI because they also start doing a host of optional tasks. But for a knowledge worker who could previously do only three things a day, a fivefold improvement can genuinely save time.
Discussing consensus and disagreement between China and the United States, Yang offered two sobering qualifications.
The first challenged the “caught-up” narrative. Chinese open-source models have improved enormously over the past six months and have indeed put pressure on several leading US companies.
But part of the apparent narrowing of the gap, he argued, reflects policy and safety considerations affecting US frontier-model releases: the strongest capabilities are not necessarily made available immediately.
He does not believe the actual gap between Chinese and US frontier models has clearly narrowed. Rather, the absolute capabilities of the models available for widespread use in China have been improving extraordinarily fast.
And that, to him, is more important. He has never cared about who scores higher on a benchmark; he cares about how many people benefit and at what scale.
A related pricing conclusion follows: as long as state-of-the-art capabilities continue to improve and have not hit a wall, frontier models will retain pricing power.
A crowd of open-source models catching up from below will not completely drive down the price of the genuine state of the art.
The second qualification concerned FDE. The term has been fashionable this year, but Yang said the work has always existed; a new name has simply made it a hot topic.
The difference from traditional customer success is this: a customer-success employee might visit a client, click through a system and configure it, but would not go back and alter the company’s own code. FDE has a feedback loop: it can change the core system directly and feed needs encountered at the customer site into product iteration.
The term is now being overused, however. Outsourcing, custom development and customer success are all being put into the FDE basket.
After those qualifications, Yang changed direction: outsourcing in the AI era really may be different.
He described the success rate of Chinese software-outsourcing projects as so low that he did not want to state the figure on the livestream. Many projects cannot pass acceptance, and payments cannot be collected.
The vicious circle is simple: the client pushes down the price, the supplier can only reduce delivery costs, the result gets worse, and the client pushes the price down again.
AI may break that deadlock. Work that once required an engineer earning 50,000 a month might now be done by an engineer earning 10,000 a month who spends another 10,000 a month on tokens, Yang suggested.
If success rates rise, the entire Chinese software industry could enter a virtuous cycle.
There is a larger opportunity beyond that. Chinese companies generally did not build strong information systems or complete their digitization in the previous era, leaving substantial work undone.
But the AI era may not require them to retrace the old route.
Yang offered a striking analogy: in the 2000s, banks gave away gifts to persuade people to get credit cards. Just when China seemed poised to catch up with developed countries where everyone had one, mobile payments appeared and changed the track entirely.
Chinese companies’ unfinished SaaS adoption might follow the same story: instead of completing that path, they could move straight from SaaS to agents.
AI itself can help companies automatically record, synchronize and process data they previously could not capture. That change is already happening, he said.
There is an indispensable condition, however: operating data and processes must exist in the digital world.
AI works in the digital world. Wherever a business lacks digitization, AI cannot help it.
If information systems are missing, build them first; once the information is organized, digitize it. That step cannot be skipped.
At the end of the livestream, I asked a question on everyone’s behalf: what should we do next?
Yang did not offer a motivational slogan. He offered a demanding test.
Suppose model capabilities make a major leap every three months—which, he said, is effectively today’s pace. Look back over the past year or two: have you updated your working methods and your understanding of AI every three months?
If not, then at some point you must have gone wrong, in his view.
He also described a personal experience: the overnight change from GPT-5.3 to 5.4.
The day before, he had been using 5.3 to write an iOS app. When 5.4 came out the next morning, he continued the same work. The model suddenly told him: “I see that your iOS device is connected to the computer. I can deploy automatically, click through tests automatically, generate logs automatically, retrieve the results, check and revise the work myself, and then release it again.”
That was the extent of the difference between two versions, he said.
That day, he sat at his computer drinking tea. Most of the work no longer required his hands.
You cannot feel that shock simply by reading the news. You have to try it yourself.
Even Yang admits he cannot try every new model. But he still tries the important ones: be hands-on; get your hands dirty.
Business owners who delegate the entire AI transformation to others and never engage with the front line themselves are unlikely to succeed, he argued.
The final slide of his in-person presentations in 2023 said: “For the next five years, cherish your time with the people you love.”
Many people took it as a joke then. Now, with a year and a half left in that five-year window, none of the convictions he has accumulated along the way has changed.
Here is a question to leave you with: will you work the same way three months from now as you do today?
If the answer is yes, that may be dangerous.
Guest
Yang Pan, founder of Bilink and former co-founder of SiliconFlow
Host
Wu Wei, founder of Unique Research
Wu Wei: Let me start directly. When did you realize that this wave of AI was genuinely different and decide to go all in?
Yang Pan: I usually put it this way: I have faith in it.
I used ChatGPT as soon as it came out. That would have been December 2022.
After trying it, I realized it was completely different from the NLP systems and other things we had used before.
Many people were still arguing about whether it actually had intelligence. My feeling at the time was that we had already broken through the barrier.
The years that followed, right up to today, have simply been the process—a question of timing. That is my first conviction.
My second conviction is that, from the moment ChatGPT was released, I believed we could reach the threshold of AGI within five years.
People may define AGI differently, of course, but I have consistently used that five-year horizon.
In my in-person talks in 2023, I was talking about five years later; in 2024, I was still talking about that five-year horizon. How much time do I have left this year? Perhaps a year and a half.
Wu Wei: What would tell you AGI has arrived? Passing the Turing test, or the appearance of some particular capability?
Yang Pan: That criterion has actually kept changing too.
Previously, I might have thought in terms of AI replacing humans in certain fields or parts of a process. Today, “replacement” no longer seems sufficient, because it can already replace people in some situations.
I now lean more toward it genuinely surpassing humans on certain problems, jobs or tasks.
I also think that moment must be an inflection point. After it, there should be a major leap, an explosion or a sharp turn; that is what would make it meaningful.
Otherwise, looking back at the curve, humanity might have passed AGI quietly without noticing.
It may already have happened in a laboratory. But from the perspective of human society as a whole, I believe there will be an inflection point everyone can feel.
Wu Wei: A little like the “singularity”: after that point, everything starts to be different.
Yang Pan: Yes. I believe that singularity will come. The difficult question is which direction things turn afterward.
Since 2023, the last slide of my in-person talks has often carried this line: “For the next five years, cherish your time with the people you love.”
Because it is genuinely hard to say what things will look like five years later.
Wu Wei: Another important conviction of yours concerns tokens. People talked about Token Maxxing, then started feeling tokens were too expensive and should be minimized. Should we maximize or minimize now?
Yang Pan: First, when different people talk about tokens, they actually mean completely different things.
Look upstream and downstream from tokens: there is computing capacity, there are models, there is inference, and finally there are the tokens we actually consume in Codex or a work agent. Every layer has its own economic and industry model.
But from an industry perspective, one very big change over the past year is that we have begun finding scenarios that genuinely consume tokens at scale.
Think about asking an agent to solve a problem in the first half of 2026, compared with asking a model the same question when DeepSeek first appeared in 2025. Token consumption behind the scenes might differ by ten times, 100 times or even 1,000 times.
That is what has changed over the past year.
As the effect continues upstream, it drives GPUs, memory and the entire infrastructure stack.
Coding agents and work agents are a particularly important reason.
When a new group first encounters them, it inevitably sees the upside first: they help me write more code, complete tasks and gain productivity leverage.
Once people see leverage, they immediately increase it.
After doing that for a while, however, they begin to see problems.
You may produce a great deal, for example, but it is not necessarily valuable. Some has no commercial value; some is simply so-called AI slop.
So I see a very clear shift this year.
After Lunar New Year, many large technology companies invited me to give internal talks. Their main concern was “how AI can improve organizational efficiency.” Half a year later, the question has become: how do we actually unlock that efficiency and those returns?
That suggests the market is beginning to cool its head.
Token Maxxing is, in a sense, a “peak of ignorance” in the technology cycle.
People are now moving from “Who burns the most tokens?” toward “What value do these tokens actually create?”
Wu Wei: If you were to summarize the most important areas of global AI consensus so far in 2026, what would come first?
Yang Pan: Model capability still comes first.
I think model capabilities have improved enormously over the past six months.
That may sound unremarkable—haven’t they been improving for years? But there has always been a very mainstream argument: have models hit a wall? Have scaling laws encountered a bottleneck?
There is another argument too: with Chinese open-source models improving so quickly, have they already caught up with the United States?
There is some truth in these arguments, but my judgment is that improvements in model capability are far from over.
Behind this latest round of improvement is something I consider important: we have obtained another body of human experiential data.
Previously, model training mainly consumed humanity’s existing knowledge, webpages, literature and information.
Then Claude Code introduced Skills. Many people understand a Skill as a tool—“it helps me do things”—but from the perspective of model development, I consider it a very meaningful invention.
After Skills appeared, human experts began contributing much of the experience they had previously kept to themselves in the form of Skills.
In a sense, humanity has begun distilling itself.
Anthropic later explained that Skills can be of different types.
One is Capability Uplift.
The model could not do something before. You give it a Skill, effectively an instruction manual or a set of special techniques. This version needs the Skill to complete the job, but the next version may have internalized the capability. Even without being given the Skill, it can do the job itself.
Another is Encoded Preference.
That has nothing to do with capability. Examples include “I want you always to follow steps 1, 2, 3, 4 and 5,” “I like this visual style,” or “Our company’s presentations all use this template.” These are preferences and standards; there is no need for a model to turn them into a general capability.
So we should not simply understand a Skill as a plug-in today.
It is a new route by which human experience becomes model capability.
I also wrote about DeepSeek Harness recently. I think things in the Harness category could bring another round of capability improvements: combine a model with a harness and let it keep training itself through approaches such as reinforcement learning.
So my judgment remains the same: model capabilities are nowhere near a wall. There is still a great deal of room ahead.
Wu Wei: Does that mean “large models will swallow everything”? First we wrote prompts, then models absorbed prompts. Later, workflows were absorbed. We put Skills into agents, and those Skills may eventually be internalized by models too.
Yang Pan: I have a fairly aggressive prediction: models may ultimately capture 80% of the economic gains in the global software or internet market.
I also think models and agents will become increasingly integrated.
Discussing models in isolation, or agents in isolation, will make less and less sense.
Wu Wei: What is the remaining 20%?
Yang Pan: That is a particularly worthwhile question.
If models and agents really take 80%, everyone else must ask: what cannot be trained into a model?
For example, we are in Hangzhou today, and there has just been heavy rain outside.
Ask a model, “What is the weather usually like in Hangzhou on August 31?” and it can estimate an answer quite close to reality from historical data.
But ask it: is it actually raining outside at this very moment?
That information cannot be trained into it in advance.
The world contains an enormous amount of data that is generated in real time, changes continually and needs processing.
Whoever controls that data may ultimately be someone models have to turn to.
Another category is capability interfaces.
For example, a model can help you do many things in the digital world. But if you say, “Order me a meal,” someone in the physical world still has to bring it to you.
Much of China’s industrial manufacturing also takes place in the physical world.
So I consider real-time data, the physical world and capability interfaces important sources of value outside models.
Wu Wei: What is the second major area of consensus?
Yang Pan: I think the second is particularly synchronized between China and the rest of the world.
Previously, agents were mainly coding agents.
In 2025, people primarily used Claude Code, Cursor and Codex to write code.
This year, though, the expansion from coding into office and knowledge work has become especially clear.
Almost every major technology company, in China and abroad, has begun committing resources in this direction.
Why does that matter?
China has programmers on the order of ten million—and that may already be a broad definition.
But knowledge workers who sit at computers in China number on the order of a hundred million.
The move from coding agents to work agents therefore represents a tenfold expansion in the potential user base.
DeepSeek’s breakthrough around Lunar New Year last year brought a huge expansion beyond the original audience.
For the first time, many people in central and other state-owned enterprises and in non-IT industries realized: there is this thing called DeepSeek; perhaps I can use AI in my office work too?
This year, their understanding has advanced another step: agents really can do work for me.
A year ago, agents might not have been good at making presentations or working in Excel. This year, they are entering those scenarios at scale.
I see that as a major consensus of 2026, one that has even changed which branch of the AI technology tree continues to grow.
But an overwhelmingly strong consensus is not necessarily entirely good.
In one talk during the first half of the year, I said that the first half had actually been rather “boring.”
Why? The consensus was too strong and too uniform.
All the major companies, talent, capital and traffic rushed in the same direction. Many smaller directions that might be equally worth exploring were drowned out.
For startups, a diverse market is actually friendlier.
Wu Wei: Looking at the evolution of agent products, I see several important milestones: Manus in March 2025, the later popularity of OpenClaw—the “little lobster”—and now work agents.
Yang Pan: I completely agree.
Before Manus appeared, there was no genuinely shared global understanding of what an agent actually was.
After Manus, a very concrete consensus formed for the first time: “This product looks like Manus, so it is an agent.”
It gave agents a tangible product form.
That gives Manus an important place in history, in my view—and globally, not just in China.
The same is true of OpenClaw.
What matters about it is not simply the number of GitHub stars.
I have always paid close attention to one question: what event drives each expansion of an industry beyond its existing circle?
Some technical advances matter only within the technical community.
What really affects an industry is ordinary users experiencing, at scale and for the first time, “So this thing really can be used like this.”
OpenClaw had that effect.
It helped many people break through a barrier for the first time: work can be delegated.
You do not necessarily have to sit at the computer watching it.
You hand over a task, it goes off and does it, then comes back with the completed work.
Today, we are accustomed to sending agents tasks through instant messaging and dispatching work remotely. OpenClaw’s influence is visible in the logic of many such products.
Wu Wei: How would you define an agent? What is its biggest difference from ChatGPT?
Yang Pan: From an ordinary user’s perspective, it is very simple: an agent is something you can delegate a task to.
ChatGPT began with conversation, then multi-turn conversation.
An agent is different.
You give it a goal. It plans, executes, checks and loops by itself, adjusts when it finds a mistake, and ultimately delivers a result.
For example, you ask, “Will it rain in Hangzhou today?” It calls an interface and gives you the answer. That is actually quite simple.
But suppose you say: retrieve all of Hangzhou’s rainfall data for this summer, organize it into a report, and make it into an attractive webpage.
A series of tasks follows.
It must retrieve data, process it, analyze it, generate a document and design an interface.
That starts to approach what an agent is.
One particularly important measure this year is that, after accepting a single delegation, an agent can keep working for longer and longer.
On the surface, that is task duration. Underneath, it is the complexity of the task it can complete.
People are the same.
Give someone the same exam with 40 minutes, four hours or four days, and the results will differ.
When a model lacks capability, asking it to work a little longer does not help.
It thinks, “That is about enough,” and hands in the paper after five minutes.
So the ability to handle long-horizon tasks is a very important indicator of intelligence.
Yang Pan: Another change is crucial.
Previously, models mainly wrote and executed programs, but were weak at operating computers and browsers.
That is where Computer Use and Browser Use come in.
Why do these two things matter?
If a task only requires generating a presentation or writing a program, a model can do it itself.
But if it needs to open a browser, log into a website, click buttons, fill things in, submit them, retrieve the result and continue, it often used to stop and wait for a person.
Once that barrier is broken, an agent’s working time can be extended substantially.
Let me give my own example.
Last year, I launched a website from scratch.
AI could write the code, but I had to buy the domain myself: open the registrar’s website, enter credit-card and other details, obtain the domain, then deploy the server.
For SEO and traffic analytics, I also had to go to Google myself, apply for Analytics and connect it afterward.
A whole series of intermediate steps had to be completed by me.
This year, that is often no longer necessary.
I tell it: I want to build a website, roughly with this name. Check whether the .io, .net or .cn domain can be registered.
It can complete the entire process itself, from code, domain and deployment through final submission to search engines.
That is significant.
Wu Wei: It is a little like the model is the brain, and Computer Use and Browser Use are the hands and feet.
Yang Pan: Exactly.
With only a brain and no hands or feet, there are still many things you cannot do.
AI also has an important characteristic that people easily overlook: its knowledge is more evenly distributed across fields than a person’s.
Why might someone find coding easy but deployment complicated?
Because they have written a lot of code and done little deployment.
AI is different: its knowledge spans a very broad range of fields.
In the first stage, we kept asking: what could conventional programs not do before that AI can now do?
In the next stage, we should start asking: what can people do, but AI is inherently better suited to doing?
For example, can a person work on a task for 72 hours without sleeping? No.
A person gets bored after an hour of repetitive work. AI does not.
A person also has at most 24 hours in a day.
But if one AI can do the job and I have enough tokens, I can theoretically have 1,000 AIs do it at the same time.
Yang Pan: There is an analogy I particularly like.
The internet solves the problem of distance in space.
To meet with an American, I do not need to take a flight; I can open Zoom.
AI, in a sense, solves the problem of time.
It can compress time, and it can expand it.
A person cannot click a mouse 1,000 times simultaneously, but 1,000 agents can.
So using AI gives us two enormous sources of leverage.
The first is borrowed intelligence.
Models possess a very broad body of human knowledge. They let me accomplish things I previously could not do.
The second is borrowed time.
If both AI and I can do something, I can delegate it and go to sleep. It has effectively “borrowed” that portion of time for me.
Wu Wei: My own experience feels like the opposite. After using AI, I have even less time.
Yang Pan: Yes, that is where another phenomenon appears.
AI suddenly lets us do many things we could not do before.
I call this “compensation.”
I used to want to paint, but I had no artistic training, so that desire remained unmet. AI can now let me paint.
I used to want to make a piece of software, but I could not code, and it was not worth hiring a programmer specifically for it. AI can now make it for me.
It is making up for things we once “wanted but could not get.”
But there is a problem: satisfying that previously unmet desire does not necessarily mean it has value.
Yang Pan: In the past, making software or a product was expensive. You had to believe it had commercial value before investing time, money and hiring people to develop it.
So “creation—distribution—commercialization” used to be largely integrated.
Now that AI has made creation very inexpensive, the three are beginning to separate.
Many people make a large number of things.
Making them feels wonderful; the dopamine is high.
But afterward, nobody uses them and they make no money.
That does not mean creation itself has no value. It means you have not completed the next two steps.
First, market it and distribute it.
Second, establish a business model so that value comes back.
For now, AI mainly accelerates the first stage: creation.
Traffic has not suddenly become infinite.
Users’ wallets have not suddenly become infinite either.
Marketing, acquiring users and getting people to pay still cost real money.
Wu Wei: Does that also explain why the path of the so-called OPC, or one-person company, is not as broad as people imagined?
Yang Pan: Yes.
The OPCs that are genuinely producing positive value tend to be concentrated in fields where an individual can also close the loop on the next two steps relatively easily.
Independent media and AI video, for example.
But developing a product on your own, then distributing, marketing and commercializing it alone, remains very difficult for most people.
Look at it another way, however, and this will also change the entire software industry.
Previously, a user with a need would search for it, try three to five programs and choose the best match.
Often, you only wanted to solve a very small problem but had to buy a very large piece of software.
The future may be different.
Much software may no longer need to be “mass-produced and mass-distributed.”
A user encounters a problem and asks an agent to generate a solution on the spot.
Once the problem is solved, that “software” might even disappear.
So I think the traditional general-purpose software market may contract, while personalized, instantly generated software becomes more common.
Wu Wei: If my computer overheats, for example, I previously had to download diagnostic software.
Yang Pan: Exactly.
Now I simply tell Codex: my computer is overheating; find out what is causing it and fix it.
I do not even need the step of “looking for cleanup software.”
Wu Wei: Taking that further, will software necessarily need a user interface in the future?
Yang Pan: My understanding of that question has changed quite a lot over the past year.
I used to advocate “generative software.”
I thought: today’s interfaces are hard-coded; future interfaces should be generated in real time according to need.
Then I realized I was still trapped by an ingrained assumption: subconsciously, I still thought software had to have an interface.
I am no longer so fixed on that.
Whether an interface is necessary ultimately depends on which form of interaction is more efficient.
For some things, two clicks in a GUI are fastest; keep using the GUI.
For others, saying one sentence is much faster than clicking through a dozen dropdown menus; use speech.
What really matters in the future may not be “what the software looks like,” but whether it has understood and fulfilled your intent.
You can use voice or text. In the future, it might even proactively recognize your context, environment and intent, then act directly.
An interface exists only when needed.
Wu Wei: That brings us to a crucial question. Many employees already use AI very well. Does that mean their company is doing a good job of AI transformation?
Yang Pan: We need to look at both sides.
If you look only at the work an employee is responsible for, using AI will usually improve efficiency.
The problem is that the efficiency dividend does not necessarily flow to the company.
Wu Wei: They might use it to slack off or work on a side business?
Yang Pan: That depends on whether the person watching is the boss or the employee. [Laughs.]
But there is a more fundamental issue.
We used to say that one person using AI might become ten times as efficient.
With two or three people, perhaps three to five times.
With five people, perhaps two to three times.
With more than ten people, the final improvement might be only 20% or 30%.
Why?
Because another name for an organization is collaboration.
When something moves from one person’s brain to another’s, it must pass through speech, writing and documents.
There will inevitably be information loss, misunderstanding and lost time in that transfer.
Thinking a problem through in your own head may be ten times or even a hundred times faster than bringing another person into a meeting to explain it.
So the more people there are, the more information crosses between brains, and the more efficiency falls.
Wu Wei: So your provocative claim can be understood this way: individuals all using AI does not mean the organization as a whole gains the same efficiency improvement.
Yang Pan: Exactly.
That forces us to rethink why we still organize companies the old way.
In the past, an individual’s knowledge and capabilities were limited, so we needed dedicated people in dedicated jobs.
Product managers understood product, developers understood development, operations staff understood operations, and marketers understood marketing.
But if one person can now cover more of the complete chain with AI’s help, why not redeploy people differently?
Yang Pan: Everyone is now exploring new ways of organizing.
One is a front-office/back-office model.
The back office builds things and tries to form a self-contained loop.
The front office handles operations, marketing and results, also forming a self-contained loop.
The two sides minimize repeated communication.
Anthropic also once shared an approach internally that divides work by roles.
For example, a Prototyper rapidly makes prototypes; a Builder actually builds the thing; a Sweeper trims away AI-generated “slop” and assorted mess to make it reliable and stable; a Grower handles growth; and a Maintainer handles long-term maintenance.
Notice that the division is not by traditional functional department. Each person is responsible for a relatively complete outcome loop.
Wu Wei: I once heard a guest say something interesting: the best collaboration is no collaboration.
Yang Pan: From the perspective of reducing information loss, yes.
But I do not think we already have “best practices for AI-native organizations.”
Not yet.
And even when things stabilize, I do not think there will be only one answer.
Companies have different businesses, and their products are at different stages.
When searching for product–market fit, for example, you should have more Prototypers, Builders and Growers.
Once the system matures, you need more Sweepers and Maintainers.
An infrastructure product called 100 million times a day and a small internal tool used by 1,000 people have completely different requirements for reliability, safety and process.
So there will not be one standard answer for AI organizations.
Yang Pan: Everyone needs to accept a basic principle: every era’s productive capabilities correspond to a set of production relationships.
Software development began with waterfall.
It later became agile.
Why did agile become possible?
Because the cost of trial and error fell.
I can make a prototype first, change it if it does not work, then change it again if necessary.
Construction, though, is the most extreme form of waterfall.
You cannot build a tower to the twentieth floor and suddenly say, “This is wrong; let’s start again.”
The cost of starting again is too high.
From day one, therefore, you must draw plans, plan the work, make Gantt charts and calculate every step.
Software used to require extensive architecture design, test cases and documentation for essentially the same reason: trial and error was expensive.
Now that AI has arrived, that cost has fallen by another order of magnitude.
If it is wrong, start again.
Methods that made sense in the past therefore do not necessarily make sense today.
What matters is understanding the relationship between productive capabilities and ways of working.
Once you understand that, you can return to your own company and look for new organizational approaches.
Wu Wei: FDE is another term people love talking about now. How do you see it?
Yang Pan: Let me pour a little cold water on FDE first.
FDE is not a completely new kind of work that suddenly appeared.
People have always done similar things. Later, it acquired the name FDE and suddenly became fashionable.
When I was a startup CTO, for example, I was doing something like FDE myself.
I would go to a customer site and use our own products and technology to help solve business problems and deliver results.
But there is a major difference from traditional customer success.
In traditional SaaS customer success, going on-site may mainly involve configuring a system and clicking a mouse.
Typically, that person would not then directly modify the core product code because of a customer’s problem.
FDE adds a loop back.
While solving problems at the customer’s site, the person can change code and the product, then feed that experience into iterations of the core product.
That is the important part.
Wu Wei: Yet much of what is called FDE today feels like IT outsourcing.
Yang Pan: Yes. As the term has expanded, the boundaries have blurred.
Customer success, outsourcing and custom development may all get put inside FDE.
But that is not the important thing.
What matters is that outsourcing in the AI era will differ from outsourcing in the software era.
Previously, developing something for a customer might genuinely require sending people to write code there for three months.
With AI today, you might discuss the requirements and have a demo right there on the spot.
That will substantially change delivery efficiency in China’s software industry.
Many software projects failed to pass acceptance in the past not necessarily because customers deliberately refused to pay, but because the work genuinely had not been completed successfully.
The budget might only cover an engineer earning 10,000 a month, while the work really required someone at the 50,000 level to do it well.
Today, that 10,000-a-month engineer, with an additional token budget, might deliver a result that previously required a 50,000-a-month engineer.
The success rate and commercial cycle of the entire software industry could then improve.
Wu Wei: But the people inside an enterprise who understand the business best are already domain experts. Why can’t they use a coding agent or work agent to build it themselves? Why do they still need FDE?
Yang Pan: In some scenarios, they may indeed not need it in the future.
But today we are still in a transitional phase.
The process from a business expert directly explaining a need to an agent reliably delivering a genuinely usable result is not yet completely mature.
For now, we still need people with technical backgrounds to help business experts bridge that gap.
That is FDE’s value at this stage.
Wu Wei: You have mentioned context repeatedly. How do you build your personal context?
Yang Pan: I take collecting my own context very seriously.
For example, I try to record nearly all my in-person presentations, important parts of conversations and my own thinking as audio, then transcribe it and gradually turn it into my own memory.
For this trip to Hangzhou, for example, I can ask an agent: look back over the past six months or year and see who mentioned meeting for a meal when I came to Hangzhou. Who have I been communicating with frequently recently, and should I see them while I am here?
Without historical context, it could not do that at all.
Wu Wei: The same applies to an enterprise.
Yang Pan: Yes.
For AI to create genuine value in a business, there are two basic prerequisites: information systems and digitization.
AI works in the digital world.
If your production and operating data have not been fully recorded there, AI cannot help.
Wherever data is missing, it cannot have an effect.
Likewise, if your operating flow, workflow and approval flow have not entered the digital world, AI does not know how to optimize them.
Your data, processes, explicit knowledge and tacit knowledge therefore ultimately need to be accessible to AI.
However powerful AI may be, without context it is water without a source.
Wu Wei: Chinese companies generally have not had such strong foundations in information systems and digitization. Do they need to go back and make up the lessons of the SaaS era?
Yang Pan: I actually think they may not need to follow the previous playbook completely.
China’s previous generation of software industries had a relatively thin foundation. Enterprise digitization and information systems were generally less mature than in the United States.
But AI may offer a new chance to change tracks.
We can compare it with mobile payments.
In the 2000s, and even the early 2010s, banks were giving away gifts aggressively to get Chinese people to take out credit cards.
Everyone in developed countries abroad had a credit card, and we seemed to be catching up.
Then mobile payments suddenly arrived.
China did not finish traveling the credit-card road; it moved directly onto another one.
AI may offer Chinese businesses a similar opportunity today.
Previously, much data could not be recorded at all.
Today, AI can process voice, video and natural language.
Things that were difficult to digitize before may be digitized directly in the AI era.
Could China’s software industry therefore bypass part of the traditional SaaS stage and move directly toward agents?
I think it could.
Wu Wei: What do you think is the biggest area of disagreement about the Chinese and US AI markets?
Yang Pan: I think China keeps running into the same issue: revenue, payment and commercialization.
The services sector accounts for a large share of the US economy. Law, finance and healthcare are enormous software markets in their own right.
China’s industrial structure is different.
China’s enterprise software market also developed a vicious circle: clients thought the software was poor, so they pushed down the price.
The more suppliers were squeezed on price, the less they could invest in good engineers, and the worse the delivered product became.
The client then concluded that it really was poor and squeezed the price again in the next round.
The consumer side is similar.
US users have a well-established subscription habit. China relied more on advertising in the past.
Tencent Video and QQ Music have helped develop subscription habits to some extent, but subscriptions still represent a relatively small share of the overall internet economy.
With AI’s arrival, Chinese users are beginning, for the first time, to pay very directly for tokens—for “intelligence.”
Whether that can change the previous commercialization structure is very much worth watching, I think.
Wu Wei: At least over the past six months, Chinese open-source models have been very strong globally.
Yang Pan: I agree.
Chinese open-source models have made enormous progress over the past six months and have put some pressure on leading US companies.
But I think everyone should remain clear-headed.
Part of the apparent narrowing of the China–US model gap today may be that some of the most advanced US capabilities are affected by policy, safety and other considerations and are not necessarily released to the market in full immediately.
To my knowledge, US frontier models themselves have not stopped improving.
So if the question is whether the real gap in frontier capabilities has already narrowed significantly, I am not that optimistic.
Another perspective matters more, however: the absolute capabilities of the models ordinary Chinese users can access today are improving very quickly.
That is enormously significant.
I have always cared less about first or second place in rankings than about how many people genuinely benefit from these models.
The overwhelming majority of ordinary people may never use many of the capabilities measured by benchmarks in their entire lives.
But Chinese open-source models allow products such as WorkBuddy to serve more people with better models at lower cost. That value is real.
There is one other point.
As long as state-of-the-art capabilities keep improving and have not hit a wall, frontier models will retain pricing power.
A crowd of open-source models catching up from below will not completely drive down the price of the genuine state of the art.
Wu Wei: Let us finish with ordinary people. I might be an employee, a manager or an entrepreneur. At this point in 2026, what should I most be doing?
Yang Pan: It is difficult to predict what will become consensus and what will move outside it.
The vaguer the prediction, the higher the chance of being right, of course. Get specific, and you are often proved wrong.
But I think there is one thing we can believe: model capabilities will continue improving rapidly.
Many people have recently asked me what deserves a long-term commitment in the AI era.
My answer is that, at this stage, “continuous, rapid change” is itself the long-term commitment.
You have to believe that it will keep changing quickly and getting better.
If models make a major leap in capability every three months, think back over the past year or two: have you revised your working methods and understanding every three months?
If not, I think you must have made mistakes at certain stages.
Wu Wei: But doesn’t that mean something you learn today gets overturned three months later?
Yang Pan: Yes.
That is why you cannot just chase surface-level products and techniques.
Products change every day. Today they have one name and tomorrow another; you will never keep up with them all.
Try instead to understand what lies underneath.
Why does organizational collaboration create loss? Why must production relationships change when productive capabilities change? Why do working methods change when the cost of trial and error falls?
Those things are relatively stable.
At the same time, you must keep practicing.
You must be hands-on.
When a new model or agent appears, you do not have to try every one. But in areas relevant to you, you must try them yourself.
If you never get hands-on and only read media reports, you will never feel the extent of the technological change.
Even working in the industry, I now find there are too many new models to try them all.
But I choose the important ones to test.
One experience I remember particularly clearly was the move from GPT-5.3 to 5.4.
The day before, I had been using 5.3 to develop an iOS app, and I had to do many things myself.
The next day, 5.4 was released, and I was still working on the same task.
It suddenly told me: “I see that your iPhone is connected to the computer. I can deploy automatically, click through tests automatically, generate logs automatically, retrieve the test results, check and revise the work myself, then release it again.”
That much had changed between just two versions.
That day, I basically sat at the computer drinking tea and watching it work.
You cannot feel that experience without getting hands-on.
That applies to individuals and business owners alike.
If an owner delegates the entire AI transformation to someone else, never tries it, stays away from the front line and has no firsthand experience, I think they are unlikely to do it well.
After these two hours, I do not actually want everyone to remember one particular tool.
I would rather you take away a way of thinking: look beneath the surface, accept change and keep practicing.
That will last longer than remembering which model ranks first today or which agent is the hottest.
Original Chinese interview by Unique Research, September 6, 2026.
The opening editorial digest and complete interview are both retained, including repeated ideas. First-person statements belong to the original host or the identified interview speaker. Estimates, predictions, product-history claims and personal anecdotes reflect their views and accounts, not independently verified findings. The original does not specify the currency in its monthly salary and token-spending examples; the amounts are retained without assigning a currency.