
Original · Unique Research · 2026-07-12
Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the SandBaseAI / David (Li Yangbing) interview narrative, the three-stage industry migration, the product and commercialization sections, and the closing outlook. All named companies, products, and figures are preserved. Founder statements and funding figures are source attributions, not independently verified findings.
AI Industry Observation
The Agent's next battlefield isn't the model — it's the Runtime.
"The real bottleneck of the Agent sector has already moved from 'can the model do it' to 'can it actually run after it's built.'"
Making an Agent demo takes a few hours, but going live means integrating seven vendors.
This isn't exaggeration; it's the daily life of an Agent Builder in 2026. Li Yangbing David (founder of SandBaseAI) said something to me I think is very true: "The biggest problem now isn't that Agent demos can't be made, but that demos are hard to turn into a truly runnable, deliverable, scalable Product."
For example: calling Claude Code or Codex to write an Agent that runs once is an afternoon's work. Open Cursor, write a few lines, connect an API, watch a beautiful chain of reasoning print in the terminal, post a screenshot to Twitter, collect a few dozen likes — beautiful.
But if you want it to keep running, handle errors, record state, execute code safely, operate the browser automatically, connect external tools, save logs, and recover when something breaks — congratulations, you need to handle a whole vendor matrix.
Model from one, sandbox from one, browser execution from one, search API from one, logs you build yourself, state you store yourself, permissions you manage yourself.
Each item alone is cheap and has mature vendors. OpenRouter does model routing, E2B does code sandboxes, Browserbase does browser automation, plus various observability and logging tools. But put seven bills together, then add the engineering time to string them together — before this Agent has served a single real user, you're already in debt for a pile of technical debt.
This sounds fragmented, but it's exactly the real daily life of most Agent Builders in 2026.
Models get stronger and calls easier, but going live with an Agent gets more painful. Claude can write you an Agent that calls APIs in three hours, but next you face: how to isolate the execution environment? How to run code safely? How to stabilize browser automation? Where to store state? How to trace when something breaks? How to tier permissions?
"David calls all these the Runtime — the Agent's runtime layer. It's unsexy, holds no launch event, and has no flashy effects in demo videos. But without it, the Agent can only ever live in a Jupyter Notebook."
In plain terms, the Agent sector's real bottleneck has already moved from "can the model do it" to "can it run after it's built."
From "Acceleration" to "Foundation": The Agent Industry Is Going Through Its Third Migration
David's background is interesting. He's long worked in AI Infra, cloud-native infrastructure, and developer platforms, done large-scale system architecture, and been on the front lines of AI Infra startups. His understanding of this industry didn't come from papers, but was ground out in building, operating, stepping in pits, and filling them.
After talking with him, a three-stage theory surfaced in my mind — the Agent industry is going through its third migration.
The first stage: everyone fights over model capability. Whoever has more parameters, higher benchmarks, can get funding. Back when GPT-4 released, the model gap was the product gap. Everyone compared MMLU, HumanEval, and various leaderboards.
The second stage: everyone fights over calling and inference efficiency. After model capabilities converged, the question became who can call models faster and cheaper. API proxies, inference optimization, model routing — OpenRouter is the signature product of this stage, with daily calls already past a hundred billion tokens. Various inference frameworks, quantization schemes, and edge-deployment tools emerged one after another.
The third stage: after OpenClaw and Hermes appeared, what's truly missing changed. Not the model, not call speed, but the execution environment, state system, tool system, and runtime control plane.
What does that mean?
When an Agent goes from "runs once and ends" to "needs to run continuously," it's no longer as simple as "model + prompt." It needs an environment to execute code, state to remember previous operations, tools to interact with the outside world, and a control plane to manage permissions and observe runtime. It needs to know "what did I just do," "if this step fails how do I roll back," "does this operation have permission."
"David's judgment is direct: the previous two generations competed on 'acceleration' — making the model called faster. The third generation competes on 'foundation' — giving Agents a stable infrastructure to run on."
That's how SandBaseAI came about. Sandbox (secure isolation) + Base (infrastructure foundation) — the name itself states its positioning: an AI Agent-native Runtime Platform.
Not building models, not building applications, but building the layer that lets Agents run. It doesn't do the Agent's business logic, but provides the complete environment for the Agent to run — model routing, sandbox execution, browser automation, state management, log observability, all packaged in one platform.
Why this timing? David believes Agent capability has crossed the "usable" threshold — Claude 3.5 Sonnet's tool use, GPT-4o's function calling, the emergence of various MCP protocols all show the Agent's core capability is ready. But capability ready doesn't mean product ready. Just as smartphone hardware was good enough in 2007, but what truly ignited the app ecosystem was iOS's Runtime and App Store — that layer of infrastructure that let apps run, distribute, and iterate continuously.
Builders Don't Lack Ideas and Tools; What They Lack Is the Layer That Glues Fragments Into a System
You may ask: do Builders really need this? There are so many tools now; can't they just piece it together themselves?
Honestly, piecing it together yourself is precisely the problem.
Today's Agent Builders — especially teams that have made a demo and are trying to move toward a product — simply don't lack ideas. Model with Claude or GPT, tools with Composio or custom-built, browser with Browserbase or Browser Use, logs with LangSmith or in-house — all have mature options.
But what they lack is the layer that "pieces these into a system."
David described a typical scenario to me: after an Agent truly goes live, it doesn't run once and end. It needs an environment, state, tools, permissions, observability, and recovery. For each item you have vendors to choose from, but piecing them into a complete chain — the engineering and maintenance in between far exceed what most small teams can bear.
For example: your Agent needs to call GPT-5 for reasoning, then run a piece of Python code in a sandbox, then open a browser to fill a form, and finally save the result.
These four steps respectively need: model API, code sandbox, browser automation, state storage. Each step has vendors, but who guarantees that if step one fails, the rest won't keep running? Who strings the logs of the four steps together for troubleshooting? Who manages the lifecycle of browser sessions and sandbox sessions?
These "middle-layer" problems won't be solved by any vendor — because that's not their business scope.
SandBaseAI's solution is: uniformly connect models and APIs; the Builder creates an Agent on the platform, configuring the system prompt, tools, MCP (Model Context Protocol), runtime environment, and execution capabilities. Then the Agent executes tasks in the Runtime — the Sandbox runs code, Browser Automation operates web pages, and Session and Trace record the whole process.
These Runtime-level needs are uniformly solved by SandBaseAI with six modules.
The input is the Builder's product goal and business flow; the output is a runnable Agent system — not a single call, but a stateful, observable, recoverable continuous-running unit.
This difference is critical. A single call is "stateless" — you send a request, the model returns a response, done. But a true Agent product is "stateful" — it needs to remember the previous round's operations, know which step it's at, and roll back or retry when the previous step fails. This state management is one of the core responsibilities of the Runtime layer.
You may ask again: how is this different from piecing together OpenRouter + E2B + Browserbase yourself?
"The difference is 'who's responsible for stringing them together.' Piece it yourself and you get a pile of independent blocks; you fill the gaps in between. What SandBaseAI wants to give is an already-built skeleton — you just put your business logic in."
The Runtime layer handles execution isolation, state flow, log chaining, error recovery — these "unsexy but fatal" infrastructure problems.
OpenRouter validated routing, E2B validated the sandbox — but the unified entry is still blank
The most interesting thing about Runtime is: point tools have already been validated by the market, but a unified entry hasn't appeared.
OpenRouter does model routing; a hundred-billion daily call volume validates "developers need a unified model-access layer." E2B does code sandboxes, raised $20 million, validating "Agents need secure code execution." Browserbase does browser automation, raised over $80 million, validating "Agents need the ability to operate web pages." Exa does AI search, validating "Agents need structured web data."
The giants are also releasing signals. Anthropic pushes MCP (Model Context Protocol), essentially defining the standard by which Agents interact with external tools. OpenAI pushes the Agent SDK and execution environments. Cloudflare is also laying out edge-side Agent runtime capability.
Everyone is moving toward the same direction — Runtime.
But here's a counterintuitive insight: the more point tools there are, the more painful it is for developers to piece them together. This isn't ecosystem richness; it's fragmentation.
Every startup that validates a point need makes developers' "piecing list" longer. Model routing chosen, sandbox chosen, browser automation chosen — but who guarantees they collaborate smoothly? Who handles cross-system state sync? Who does unified logging and observability?
When E2B upgrades its API version, do you rewrite your glue code? When Browserbase changes its session management, do your Agent's logic need to adjust?
These "assembly costs" don't appear on any vendor's pricing page, but they're real and grow exponentially with your Agent's complexity.
"David's judgment is clear: point Agent Infra has been validated, but a unified Agent Runtime entry hasn't appeared yet."
SandBaseAI's benchmark isn't any one company. The reference is more like the infrastructure layer of the Agent era — both providing the underlying execution and scheduling capability and the unified access to upper-layer tools and models.
In other words, it's not competing with OpenRouter or E2B, but trying to become "the layer above all of these." Just as Cloudflare doesn't produce CDN content sources but carries the complete content-delivery chain; SandBaseAI doesn't produce models and tools, but carries the complete lifecycle of Agent operation.
This positioning has a big advantage: it doesn't depend on any one model or technical route. Whether GPT-5 or Claude 4 wins, whether Python or JavaScript becomes the mainstream Agent language, as long as Agents need a runtime environment, state management, and execution isolation — the Runtime layer has its value. This "technical neutrality" is the hardest capital of an infrastructure company.
Commercialization Isn't Selling Models, but Selling the Agent's "Running Bill"
Speaking of business model, you may already be thinking: how does this thing make money?
First, a trap. Only forwarding model APIs is a dead end. OpenRouter has already run this model — low margin, high traffic, relying on scale. Latecomers competing on price in this sector have no chance.
SandBaseAI's revenue path is completely different: around infrastructure consumption generated by Agent execution.
In the short term, revenue comes from five segments: Tools (tool calls), API/Runtime (runtime calls), Models (model routing), Sandbox (sandbox execution), Browser (browser automation). Each is a "pay for what you use" consumption model. Your Agent calls a tool once and a fee accrues; running a sandbox for a second accrues a second's fee; opening a browser session accrues a session's fee.
This is similar to model-API billing logic, but covers the full chain of Agent runtime — not just "calling the model," but "letting the Agent run a complete task."
Long term, the space is bigger: execution environments, state systems, observability systems, GPU scheduling, deployment — as Agents move from "experiment" to "production system," this infrastructure consumption only grows.
Model call fees, tool call fees, browser execution fees, sandbox execution fees, storage and state fees, compute fees, logging and observability fees, and future possible GPU scheduling, website deployment, DNS, and security-policy fees — the longer the Agent's execution chain, the more billing points the Runtime platform has.
The early paying-user profile is clear: developers and teams that have already pieced together a round of infrastructure themselves and truly hit the pain point. These people best know the pain of "seven vendors" and are most willing to pay for a unified platform. They aren't here to try something new; they're here to solve a problem.
How to calculate migration cost? Assume a five-person team, each at $15,000/month salary, each person spending 5 hours a week maintaining seven infrastructure stacks — that's 150 engineering hours a month, over $10,000 in cost.
This doesn't even include the "psychological cost" of cross-system troubleshooting — those late-night debug moments of "is it the sandbox or the browser?" If SandBaseAI can compress this time to near zero, charging a few hundred to a few thousand dollars a month, the ROI is very positive for the team.
How to calculate ROI? Not the vague "saved some onboarding time." The real value is reducing a lot of repetitive engineering and improving Agent delivery speed, observability, and stability. A team switching from "piecing blocks" to "using the platform" saves not a few hours, but the engineering drain of continuously maintaining seven systems — and the heart-stopping moments every time a vendor upgrades.
"Don't only ask 'what technology do I have'; ask 'who uses this infrastructure every day.'"
High-frequency use + gradually rising replacement cost + carrying the business's operating chain — meet these three conditions and the infrastructure's commercial space holds. You're not a one-time tool; you're the foundation the business runs on. The foundation's value is far above the tool's price.
AI-native Companies: An Agent Can Do Many Things, But Judgment Can't Be Outsourced
After the product, I asked David a more personal question: how does your team use Agents?
His answer is down-to-earth — not "hand everything to AI," but redrawing the boundary between people and Agents.
What Agents can do: material organization, code generation, testing, documentation, research, content first drafts. These are tasks with clear input/output that can be iteratively verified; Agent efficiency far exceeds humans. An engineer using Cursor + Claude writes code 2-3x faster; using an Agent for research and preliminary organization saves a lot of repetitive labor.
What humans must do: product direction, client trade-offs, business model, fundraising pace, organizational culture. These have no standard answer, no clear "right" or "wrong," and need judgment, trade-offs, and decisions under incomplete information.
"An Agent can help you do many things, but a founder can't outsource judgment."
Team composition matters too. SandBaseAI runs a lean route — core capability concentrated in three areas.
Infrastructure engineering — able to build a Runtime, understand distributed systems, secure isolation, resource scheduling. This is the chassis; without it the Runtime can't run.
Product abstraction — able to wrap complex infrastructure into usable interfaces, used by Builders without friction. In plain terms: "keep the trouble for yourself, give the simplicity to users." For example, wrapping the dirty work of sandbox configuration, browser session management, and state persistence into an interface Builders call in a few lines — Builders don't need to understand the underlying implementation, only that "my Agent runs safely."
Developer relations — able to make Builders understand and trust the platform, building community and ecosystem. This isn't traditional "marketing," but building a Discord community, writing docs, making sample projects, so Builders spread it on their own. When a Builder runs his first Agent on SandBaseAI, will he recommend it to another Builder? This word-of-mouth effect is the healthiest growth engine for an infrastructure product.
A large, comprehensive team isn't right for this stage. The advantage of an AI-native company is using Agents to amplify each person's output, not stacking heads. Three engineers amplified by Agents may beat ten traditionally staffed developers.
This way of organizing in turn validates SandBaseAI's own product logic: if an Agent Runtime lets a three-person team produce a ten-person team's output, then the Runtime itself is valuable — it isn't selling a tool, but "human leverage."
This logic of "letting people worry less about infrastructure" has actually been validated in every technology transition.
The AWS Moment of the Agent Era
In the website era there's Cloudflare, packaging CDN, DNS, and security into an infrastructure layer developers can use on the go. In the payments era there's Stripe, abstracting complex payment chains into a few lines of code. In frontend deployment there's Vercel, making build, preview, and launch second-level operations.
Every technology-paradigm shift produces a new infrastructure entry point. What do these entries have in common? They don't directly participate in business logic, but carry the business's operation. They let developers focus on "what to do," not "how to run it."
The Agent era won't be the exception.
SandBaseAI's long-term path is three steps.
Short term, build the Agent Runtime platform — let Builders run, solve the pain of vendor fragmentation. This is the most painful point right now, and the market SandBaseAI is entering.
Mid term, build the Agent application-generation layer — Builders only describe product goals and business flows, and the Runtime auto-generates runnable websites, apps, test links, and preview environments. An application layer grows above the Runtime, evolving from "help you run" to "help you build."
Long term, become the infrastructure entry point of the Agent era — carrying DNS resolution, GPU scheduling, website deployment, execution environments, security policy, and observability control plane — the water, electricity, and gas of the Agent era. No matter what upper-layer apps look like, the underlying runtime can't bypass this layer.
This path is long, but the start is clear.
Stripped down, today's Agent market is much like cloud computing around 2010 — everyone knows it's coming, but the true infrastructure layer is still in a melee. AWS wasn't the first to do cloud, but it won because it let developers "worry less about infrastructure." The AWS moment of the Agent era may be in the next two to three years.
The Agent's demo era is noisy enough. What truly decides the winner next is who can make Agents run stably, deliver continuously, and serve real users at scale.
"That runway is still empty."