Original · Unique Research · 2026-06-28
Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the five panelist self-introductions, the three breakdown sections, the closing outlook, and the complete more-dialogue-details transcript. All named companies, products, and people are preserved. Panelist statements are source attributions, not independently verified findings.
AI Industry Observation
How exactly should Multi-Agent be done?
A group of practitioners argued for an afternoon; here's the real substance.
"
"Actually, I don't really believe in Multi-Agent."
This was said by one of the panelists present — the kind of person who works with underlying models every day and has heard "Agents will change everything" a hundred times. A shovel seller saying first that the shovel might not work — that kind of candor is rare in today's AI circle.
But it was precisely because of this sentence that the whole conversation didn't feel like a waste of time.
Not the "AI transforms every industry" forum tone, nor a read-from-the-slide fundraising roadshow. Five people sat and talked for an afternoon on one topic only: Multi-Agent — when can this thing actually be put to use?
The people who came were interesting, from five completely different sectors:
Lux Li (李粒) is at TiDB, building infrastructure for the world's top AI companies. The Multi-Agent he sees is from a database perspective — massive concurrency, state management, data consistency; these are the real pains.
Jianfeng Lu (陆剑锋) is at WIZ.AI, doing voice customer-service AI for Southeast Asia and Latin America. His scenario is blunt and simple: make a phone call; if you don't respond to the customer within two seconds, they hang up. No elegant architecture stands up to latency.
Elliot Chen (陈亦凡) runs EverMind, building Agent memory infrastructure. This sounds abstract, but it solves one core problem — an Agent starts working and forgets what it was doing; how do you break that?
Benson Xu (徐安邦)'s Loova AI makes video-generation Agents, with an extremely long pipeline — script, materials, editing, review — any step can fail.
Yi Shi (施燚) is at DeiNai, doing influencer-marketing AI Agents. His Agent has to negotiate price, contract, and schedule with influencers — real people on the other side, and the Agent has to game like a human.
The five's combined hands-on experience is worth more than any benchmark. They aren't discussing "can Multi-Agent be done," but reviewing "I've already done Multi-Agent; which pitfalls are the real ones."
Breakdown ①: When Do You Actually Need Multi-Agent?
After all the talk, one judgment criterion slowly surfaced.
Is there "role conflict" in your business?
If a task has one consistent goal and coherent logic from start to finish, a single Agent will probably handle it. But if you find different roles, different goals, even clashing forces in the business — then Multi-Agent isn't showing off, it's essential.
How to judge concretely? Four rulers.
Ruler One: Role Conflict (Lux Li, TiDB)
Lux Li drew an analogy: building an Agent system is much like building a company.
In a company, the PM wants to perfect the product, the programmer wants to deliver fast, QA wants to stop you from launching. Three roles, three goals, in conflict.
"When a task needs different Roles with different Goals, even conflicting Goals, to argue against each other — that's when you need Multi-Agent."
A single Agent is like an all-round employee; no matter how smart, you can't argue with yourself. Real decisions are ground out in conflict. Multi-Agent simulates the tension of "each fighting for itself yet forced to collaborate." Many companies jump straight to Multi-Agent architecture, when the business has no role conflict at all — pure self-drama.
Ruler Two: Capability Tiering (Jianfeng Lu, WIZ.AI)
Jianfeng Lu's voice-service scenario is very concrete. The same system needs a sweet voice for marketing and a stern tone for debt collection. Switch the prompt? No. The reasoning ability and response speed each task needs are completely different. Marketing can chat slowly; debt collection must be fast, accurate, and ruthless — if there's no response within two seconds, the customer hangs up.
"IVR systems press 1, press 2; task planning, FAQ, data query — these are already different Agents."
Making the same Agent handle high-concurrency queries, deep-reasoning complaint analysis, and voice-emotion control at once is forcing one person to be schizophrenic. Stuffing tasks with wildly different capability needs into one Agent isn't elegant architecture; it's laziness.
Ruler Three: Context Management (Benson Xu, Loova AI)
Benson Xu's angle is very technical. His video-generation pipeline is long — script, materials, editing, review, each link depends on the last. Stripped to one sentence: the collaboration Tool Calling can't handle is what needs Multi-Agent.
The key is how context flows. If throwing the full context to a Tool gets it done, a Tool is enough. But if some stages need style consistency and others need loops and backtracking, you have to give each sub-Agent its own context space. Dump everything in and the model gets dizzy; cut it too fine and you lose coherence. Multi-Agent is, in the end, a context-governance strategy, not stacking headcount.
Ruler Four: Adversarial Game (Yi Shi, DeiNai)
Yi Shi's scenario is the most human — influencer marketing. His Agent has to find influencers, negotiate price, contract, and schedule. Across the table are real influencers with emotions, preferences, and negotiation tactics.
"Simple tasks are fully handled by a single Agent; complex ones that need confrontation and gaming need multiple Agents."
This is interesting. One Agent playing both "merchant rep" and "understand the influencer's psychology" clashes. The merchant wants to lower price, the influencer wants to raise it — a zero-sum game. Better split into two: one argues the merchant's case, one understands the influencer's needs for a win-win, and let them argue to a result.
This aligns with what Lux Li said about enterprise architecture. Where there's gaming, Multi-Agent is actually working.
My understanding is this: businesses running in a straight line should first solidify a single Agent. Only when there's real role conflict, capability tiers that differ too much, context to isolate, or zero-sum confrontation do you bring in Multi-Agent.
Breakdown ②: Model Routing — Engineering Practice to Switch Models Without Breaking
Want a single Agent to switch models? Easy. Run a benchmark; if metrics don't drop, launch.
In a Multi-Agent system, this instantly becomes a nightmare. Five Agents, three models — who switches and who doesn't? A upgrades and B crashes; the knock-on effect is dominoes. Lux Li says that when they serve top clients at TiDB, model routing is the biggest engineering headache, "no standard answer, all trade-offs."
Headache aside, practitioners always have a way. I've organized four moves.
Strategy One: Treat the Regression Test Set as a Core Asset
Lux Li's approach is very engineering. Tasks are tiered first — check the weather, write an email, cheap models handle it; critical reasoning, customer-facing tasks, bring out the top closed-source model. Money is spent where the blade cuts, but how to sharpen the blade needs a standard.
Their most valuable asset isn't the model, but the regression test set accumulated internally. When a new model comes out, run the test set first; only if it passes does it go to production. Sounds like basic practice? Few teams actually do it. Most companies switch models by gut feel — "hmm, GPT-4o seems a bit faster than GPT-4" — then ship.
A tougher move is to actively make a mess. Fork multiple workspaces for the same goal, run different models in parallel, do ablation experiments. Whichever path is good, deepen it further. For uncertain things, just let them happen in parallel.
Strategy Two: The Hard Constraint of Three E's
Jianfeng Lu offered a clean framework — Effectiveness, Efficiency, Economics. Three E's; you can satisfy at most two at once.
In voice scenarios, latency is life and death. Over two seconds and the customer hangs up, no matter how well you answered. So WIZ.AI chooses models by latency first, then cost, and only last the effectiveness ceiling. Zhipu's models are cost-effective, so daily tasks use them. But high-difficulty tasks, like multi-turn emotional reassurance, still require bringing out the more expensive ones.
Every business has its own hard constraint. Voice is latency, finance is accuracy, advertising is cost. Routing strategy isn't copied from others, but first recognizing what you can't afford to lose.
Strategy Three: Let Users Vote
Benson Xu's route is more down-to-earth. Loova AI's system always runs two models — the most cost-effective as default, the best-performing as backup. Users can switch themselves, essentially voting with their feet.
But that's only the surface. The real flywheel is in user data. Which videos got used, which ad spend ROI was high, which material went viral on social — this data feeds back into the model for post-training, and routing gets smarter the more it runs. They even have people dedicated to monitoring whether videos are actually used — not reading system logs, but whether clients actually posted the generated material to their Moments.
Data feedback looks like a technical architecture on the surface, but actually tests organizational habits. Only teams that reach this layer get model routing that evolves on its own.
Strategy Four: If All Else Fails, Train Your Own
Yi Shi's solution is the most counterintuitive. The general large model can't make influencer-negotiation decisions well? Then don't let it decide.
DeiNai took four or five years of operational data, influencer communication records, and deal history and trained a vertical private model. Negotiation strategy, price anchors, influencer psychological profiles — these core decisions are made by the private model. GPT-4, Claude and other general models only do the periphery — polishing copy, generating emails, organizing notes.
Industry know-how becomes the moat; the general model becomes a senior secretary. This division is elegant — the closer to the business core, the more you need a private model; the closer to general capability, the more you can use the big players'.
After four strategies, I have one feeling: model routing has no silver bullet. Some guard with test sets, some lock with hard constraints, some let users vote, some just build their own model. The difference isn't who's smarter, but who knows their own business constraints better.
Breakdown ③: Memory Systems — Making an Agent Truly "Have Memory"
Before talking about memory, I want to burst an illusion.
Many people think an Agent's Memory is just "save the conversation history, continue next time." If it were that simple, Elliot Chen wouldn't run EverMind. The practitioners' consensus is clear — memory isn't generic storage, but a layered architecture tightly bound to the business.
The memory portion of this conversation carried the most information. Five angles, pieced into a complete picture.
Memory Must Be Layered — Casual Chat and Work Collaboration Are Two Different Things
Lux Li first gave a basic framework: is your Agent chatting with you for fun, or working for you? These two memories are completely different.
A casual-chat Agent only needs to remember what style and topics the user likes. A work Agent? It must distinguish memory by Project and support organization-level sharing. A top client raised a demanding requirement — memory must pass down layer by layer like an org chart. Details of how a junior employee talks with a client should consolidate into decision inputs for middle managers; middle managers' judgments should in turn help senior leaders decide.
Behind this is a huge engineering challenge. TiDB has done a lot of multi-tenant isolation work. "At 100,000+ users, privacy incidents are a disaster zone; one data cross-talk is a headline."
What about expired memory?
Jianfeng Lu's scenario is very concrete. A company promotion policy released in March expired in April. But a client comes back in June and asks: "Can you still give that March deal?"
If the system only keeps the latest policy, the service Agent says "we don't have that campaign," and the client explodes.
WIZ.AI's approach is hot-cold separation. Currently valid promotion policies sit in the system as hot data, queried in real time. Expired policies go to cold storage; when a client needs to trace back, an MCP tool fetches them. Time-sensitive data and historical traceability — neither can be missing.
"Born-with-it feel" — good memory makes people think it's innate
Elliot Chen's word made me laugh, but on reflection it's spot on.
He says memory must have a "born-with-it feel" — private, human, where one sentence and the other party immediately understands. Not the mechanical "you queried X last time," but "I know what you prefer; I anticipate what you might want."
How to achieve it? Preference layering plus a Forecast function. The Agent doesn't only remember what you said, but predicts what you might need next. From passive answering to proactive service, the difference is all in memory.
More aggressive is "Dreaming" — when the system is offline, the Agent reviews all of the day's execution records and evolves implicitly. What mistakes it made negotiating with people by day, which influencer's price was negotiated too high, which video went viral — at night, while the Agent "sleeps," it digests this information and wakes up smarter the next day.
LlamaIndex founder Jerry Liu said something Elliot Chen strongly agrees with: "Stop staring at the prompt; do Loop Engineering." Tuning prompts is indeed small craft; memory and feedback loops are the real work.
Don't say "Memory" to clients; say "user asset"
Benson Xu's business nose is sharp. When talking with clients, he doesn't say "Memory system," he says "user asset exclusive to your enterprise."
Three layers of asset — brand personalization data, ad-effect data feedback, user Workflow. Once Memory becomes an asset, the payer shifts from the tech team to the CEO.
Auditing Agent: smarter the longer it's used
Yi Shi's approach is the most pragmatic. Everything AI communicates with influencers — chat, price negotiation, images, video — is all stored. But not the raw record; an auditing Agent cuts redundancy and keeps only the core structured content.
This auditing Agent is itself part of the Multi-Agent system. The more it runs, the better it knows what to keep and what to discard; the whole system's memory quality rises over time.
Before building a memory system, ask yourself three questions: what to store, how to store, who can access it. What to store is a business question; how to store is an engineering question; who can access is a security and privacy question. Get all three clear, and memory becomes the Agent's moat, not technical debt dragging it down.
Elliot Chen said something that stuck with me — memory determines the likes and prejudices an Agent shows. What you store and how you organize it ultimately reflects in the Agent's behavior. It's a bit like raising a child; the environment you give, that's the kind of person he becomes.
Breaking Through: The Most Anticipated Breakthrough of This Year
Near the end, the host switched to a lighter topic — what technical breakthrough are you most looking forward to this year?
Five people, five answers. Put together, they're exactly a complete layered diagram.
At the bottom is Reliability, as Jianfeng Lu says. Anyone who's done enterprise voice service has a physical reaction to the word "unreliable." "Conversation quality, output results, remaining reliable after version iteration — this is the lifeline." If the foundation isn't solid, nothing above it works.
Up is Stability, from Lux Li. From demo to production-ready, there's a whole Pacific in between. Privacy isolation, scalability, multi-tenant no cross-talk — "top clients choose TiDB, and the core metric is this." Reliability is one point's performance; stability is the whole system's confidence.
Up again is Memory freedom, from Elliot Chen. This word has imagery. Cross-session memory continuity, cross-Agent identity consistency, data truly belonging to the user rather than locked in some platform's black box. "I want the Agent to have a 'born-with-it feel' — it's born yours, and it recognizes you wherever you take it."
Beyond memory is Self-evolve, which Benson Xu hopes for. The video-generation pipeline is too long; manual tuning will hit a ceiling one day. "The real breakthrough is building a self-evolving loop. User data comes back, the system iterates automatically, running smarter and smarter." This echoes the Dreaming Elliot Chen mentioned — the Agent is no longer just a tuned tool, but can grow its own ability.
At the top is Yi Shi's Context Management plus Workflow orchestration. In vertical-industry deployment, the algorithm is no longer the biggest bottleneck; engineering is. "How you use the toolchain to manage context well and string workflows smoothly determines whether a Multi-Agent system is a demo or a real business."
Five answers, five layers. Reliability is the foundation, stability is the load-bearing wall, Memory freedom is the feeling of living in it, self-evolve is the house repairing itself, and Context plus Workflow orchestration is the blueprint that makes all this run in the real world.
Missing any layer, and Multi-Agent can only be a pretty model in the lab.
A Bigger Lens
Finally, some big-picture talk.
Multi-Agent isn't a technical-architecture problem. Or rather, it isn't first a technical problem.
What you're building is far more complex than a more complex AI system — you're building a mini company. This company has different departments, different KPIs, conflict and coordination. The PM wants the product perfected, the programmer wants to deliver fast, QA wants to stop you from launching. Three roles, three goals, arguing, finally arguing out a result.
This is what Multi-Agent is actually doing.
Someone at the event asked: is Multi-Agent hype? Is it worth going all in?
The question itself is wrong. It's not either/or.
What you really need to think about is: does your business have scenarios needing multiple roles to game and collaborate? Is there role conflict? Capability tiering? Context to isolate? Zero-sum confrontation?
If yes, Multi-Agent is no longer a multiple-choice question, but a required one.
A single Agent is an employee; Multi-Agent is an organizational structure.
The core isn't whether AI is smart enough, but whether your business is complex enough.
Go back and look at what you're holding — if it just runs in a straight line and one smart employee can handle it, then don't rush into Multi-Agent. But if your business already has several departments bickering, arguing, and gaming, congratulations — your Multi-Agent architecture diagram may already be drawn in your org chart.
More Conversation Details
Panelists: Lux Li (李粒) (PingCAP TiDB APAC AI Business & Product Leader); Jianfeng Lu (陆剑锋) (WIZ.AI CEO); Elliot Chen (陈亦凡) (EverMind open-source ecosystem lead); Benson Xu (徐安邦) (Loova AI Founder); Yi Shi (施燚) (DeiNai CEO)
Panelist self-introductions and an initial look at Multi-Agent
Host: If you want to add my contact, X (formerly Twitter) is fastest; I often post on X about payments, large models, and graph models.
Since we're a listed company today, I can't share too much internal information, which is why I especially like being the host. Today's topic I think will be fairly hardcore — Multi-Agent. Because we're not talking about Agents as a macro concept, but cutting into details. Including myself, I'm actually not a very staunch Multi-Agent believer. Because as a model vendor, your single-agent model needs to be very smart. So today I'm really exploring this with everyone.
First, please each briefly introduce your company and your work.
Lux Li: Hi, I'm Lux Li from TiDB. Many of you may know TiDB; we started doing overseas business very early. You may know us more for providing database services to big companies, but in fact what you may not know is that we've started providing foundational software services to the world's best AI products or models.
Including Singapore's largest AI vendor, and those who haven't appeared much recently, including the OpenClaw printed on my shirt, are actually our clients. So we now provide a lot of infrastructure support for all Agent mechanisms and AI assets, including databases, Vector Services, and so on. If you're interested later, you can come talk to me.
Jianfeng Lu: Hi, I'm Jianfeng Lu, founder and CEO of WIZ.AI. WIZ.AI mainly provides customer-service and customer-engagement AI technical support for large enterprises in Southeast Asia, Latin America, and other countries. For example, in Malaysia we provide a voice dialogue system that mixes and freely switches Malay, Chinese, and English.
Another important thing is that we mainly support language dialogue on telephone systems, so we need to handle a lot of low-quality audio. Zixuan just said he doesn't lean toward supporting Multi-Agent, but I'd say the opposite — supporting Multi-Agent might be better, and everyone's opportunity will be greater. Thank you.
Elliot Chen: Hi, I'm Elliot (Chen Yifan) from Shengda EverMind. We do long-term Agent research, and I'm also active on X (Twitter); if you're often in the Chinese circle you should have come across me.
First, we do Agent Memory Infra, sustainable and evolvable Agent memory. Next, we also have a very cool C-end product coming; we've actually already released one called Everme, an integrable product. Next there's an Easter egg — we'll release our own self-evolvable Agent. I'll keep the name a surprise; you can find me at the table outside to see some interesting things.
In response to what the first two just said about Multi-Agent: how to use Multi-Agent well? The memory link is very, very critical, and that's exactly what we do.
Benson Xu: Hi, I'm Benson (Xu Anbang), founder of Loova AI. What we're doing is actually a Video Agent, because our main focus is on two scenarios: marketing video and some narrative video.
We've seen this year that video models and Agents are developing very fast, so we think the whole Agent will reshape the previous entire video production chain; that's why we're making a new Video Agent entrance, which is what we're doing. We mainly target the European and American markets.
By the way, in response to Multi-Agent: in many scenarios we won't use Multi-Agent, but when certain Context can't be passed well to a specific role's Agent, we'll use it. So it's not that every scenario needs Multi-Agent, but in some scenarios needing more, more specific Context, we'll need it; we can expand later.
Yi Shi: Hi, I'm Yi Shi, founder and CEO of DeiNai. We're a company doing Creator Marketing AI Agents. Our product lets you input your needs and, through 350 million influencers globally (such as YouTube, TikTok), help you discover and reach them.
Once found, we can see each influencer's data and judge whether to work with them. For example, if you do hardware, or AI Agents, and you want to go overseas and do marketing, and want to work with a certain influencer, click to start the project and our AI will negotiate for him. From the first email or WhatsApp communication to all the underlying details afterward, our AI does 80% of the work, then a human confirms. Just like using Copilot — write a prompt, then decide whether to do it that way. So we mainly solve the efficiency problem: what used to take 10 operators may now only need 2.
At the same time, we've also uploaded our DeiNai Creator Skills to platforms and code like OpenClaw; you can download and use them, seamlessly integrating into your Facebook and Slack. That's roughly the product.
Core Discussion One: When Does a System Need to Move from Single-Agent to Multi-Agent?
Host: Benson just said we could expand. To avoid forgetting, I'll expand right here. I want to ask all panelists: for you, how do you judge when a system needs Multi-Agent? Because besides Multi-Agent, we also have Skills, and environments — how do you judge this?
Lux Li: We're a technology software vendor, so we've seen many of the very top AI companies, model companies, and Agent companies, and also top vertical industry companies. After watching and discussing with everyone, we found one essence: building an Agent system is essentially like building a company.
When do you need Multi-Agent? It's actually when your task needs Roles with different Goals to argue — at that point you may need Multi-Agent, because these Agents need to make decisions with different purposes, possibly even conflicting purposes.
Take the simplest software development, since most here are technical. The PM (product manager) role's PRD/Spec task is to think the matter through as thoroughly as possible; the role actually responsible for Coding thinks about how to finish it ASAP to meet user delivery; the QA/Test role actually wants to block it, not let it launch casually. Each role's purpose is conflicting; in this case it's not suitable for one Agent alone, and a Multi-Agent architecture is needed. Other vertical industries also have this kind of conflict and gaming.
Jianfeng Lu: Our application scenarios are clear. For example, if I do telemarketing, first, the phone must have a fairly sweet female voice to make marketing conversion easier. But if I help a credit-card company collect bills, the voice needs to be sterner. From a Persona perspective, I need different Agents to complete these different tasks.
Second, as the TiDB colleague also mentioned, from the perspective of capability and compute. Some tasks need very high reasoning ability and high intelligence, but on others, only some very routine, simple processing is needed. So in the process, the Context they need differs; the second point is that their evaluation criteria also differ.
We provide services to many banks or Telcos (telecom operators), whose dialogue-quality requirements are very high. Each company and each role's evaluation standard may differ, so based on these I need to distribute it into different Agents so the overall effect is best. For example, the simplest bank IVR system (interactive voice response): press 1, press 2; some need task planning, some only simple FAQ, some just connect to the back-end system to query some data. So in such an AI system I may need to deploy three different Agents to complete related tasks.
Elliot Chen: This question, in my view, combining my own understanding and practice, can basically be seen from three angles:
First is "role-ification": suppose I get a complex task; can I split it and let different roles handle it? For example, an exploration Agent explores, a test Agent tests, an analysis Agent analyzes, and finally an Agent verifies.
Second is "persistence": suppose one of my Agents suddenly can't finish the task for some reason; will another Agent help complete it instead? Normally you can Pass the task to Agent B; that's persistence.
Third is "verification and review": how do I verify my output? Ultimately through what method do I verify success, and how do I review to do better next time?
Actually, in my view, when most people use so-called Multi-Agent now, they haven't reached true multi-agent self-collaboration, but are in a state of "Multi-Agent Usage." That is, in OpenClaw or code you set up several Projects, manually switching Agents in each Project, or using a Browser Agent to run some light tasks. Truly realizing Multi-Agent smoothly, complementarily and self-evolvingly may only have a chance this year. As LlamaIndex's founder said a couple days ago: stop staring only at the prompt; now you need to provide something called Loop Engineering. And to do this well, knowledge bases and memory are essential.
Benson Xu: The previous panelists have explained concepts and role division very clearly. I can share a few core judgment criteria from our practice: do we use simple Tool Calling, or make it an Agent?
First is Context management judgment. Do we input the whole Agent's context to the Tool so it can complete the task, or do we condense it into Context for specific sub-Agents? Simply, take our creative script for video as an example. For the creative-script step, is it better to throw in the whole context to execute, or only feed it the structured Context related to the creative script? If we find only giving it the creative script's Structure Context works, we make it a separate Sub-Agent.
Second is whether the task keeps running. For example, I have a background monitor watching how the ROI of the materials I've placed performs; when I need a continuously monitoring Agent, I make a separate Sub-Agent.
The third scenario is collaboration and backtracking between Agent and Agent. For example, we sometimes use a creative-script Agent to write different storyboards, but when this storyboard goes to some video models or image models, it may not generate, or the result is poor. It needs to go back and talk with the script Agent to regenerate. When there's this bidirectional collaboration and Loop, we actually also turn it into a separate Agent. Otherwise, if it's just pure Tool Calling, you actually can't do this kind of collaboration.
Yi Shi: For simple tasks now, a single Agent is fully fine. Complex ones that may need confrontation and gaming need multiple Agents. Take our Creator Marketing AI Agent: in our system, our AI Agent must both communicate with influencers, simulate influencer psychology to negotiate, and represent the merchant's interests to game, negotiate price, say how we sign the contract, this price is not OK. So at this point we use multiple Agents to game and confront. So overall, complex, multi-task, including high-concurrency tasks, we feel are more suited to a multi-Agent architecture.
Core Discussion Two: Model Routing and Uncertainty Control
Host: Then let's talk about a later topic: Model Routing. For Multi-Agent, one point I care most about is when the underlying large model upgrades (for example from 4.7 to 4.8), it releases a bunch of benchmark leaderboards, but these may have nothing to do with big companies' real production environments. If you're a Single Agent it's fairly simple — run all your benchmarks and you know the impact. But in a multi-Agent system, who switches the model and who doesn't is a very complex thing. How do you handle this?
Lux Li: We did run into this problem when developing Agents, and it's very painful. Our company internally does development and business processing with multiple models, not a single model. We actually look at multi-model routing from two angles:
First, task-tiering constraints. We tier all tasks the models process, with what level of model handling what kind of task. Simpler tasks can use lower-level, lower-cost models; very complex, critical tasks may bring out the top closed-source large model.
Second, standard constraints and assetized test sets. When we provide AI products externally, the output standard agreed with clients can't degrade because our back-end model changes. So we set key Checkpoint specifications to constrain model output. When the same model upgrades and switches, testing is unavoidable. Every company that seriously does Agents has massive internal test sets; this regression test set is one of our most valuable assets. Only when a new model passes heavily in regression testing do we switch it on.
Besides the constraint part, the second part is actively embracing routing uncertainty. For example, when doing exploratory tasks, we do breadth-first exploration. We Fork the same Workspace many times toward the same goal but with different Models, even the same Model but slightly tuned Prompt to run directly. In this case we're actively using routing uncertainty, finding better paths within the uncertainty, doing an ablation experiment, then deep optimization. This approach is a very effective behavior internally.
Jianfeng Lu: In model routing, we mainly consider three core dimensions:
Effectiveness: can this model help me do this thing well? If the quality isn't good enough, I definitely can't use this model.
Efficiency: this is critical for us handling phone voice dialogue. After the client finishes a sentence, I must respond and answer within one to two seconds (max two seconds). If the model can't manage this time, the client hangs up, so Efficiency is very important to us.
Economics: the underlying prices of different models differ greatly. For example, some Zhipu models are very cost-effective, but for certain specific high-difficulty tasks I may need others. So our model-routing choice is basically balancing these three factors.
Elliot Chen: We do memory, which isn't most directly physically related to routing, but our underlying layer can solve this well. Many people in many scenarios decide to use RAG, but actually "memory" can often replace it. Our memory Infra happens to have a Knowledge Wiki feature; with this, the different abilities of different models, after using our recall Wiki, improve very strongly.
It's a bit like last year's "hundred-model war," where everyone saw which model had different advantages for certain things. But in vertical Agents, the latest model at the legal or medical level isn't proven the best. Because models care about similarity and relevance; if you use some mechanism to strongly tag and more accurately recall, you can solve the routing problem caused by the large model's own ability changes.
Benson Xu: On model routing and multimodal selection, our product has a voice, because we deal with video, image, and large language models. When we select and control models, one core point is that the team first does rigorous evaluation. Second, we use users' usage process and feedback data for post-training; these two are actually necessary.
After testing makes these models reach a usable standard, we often retain two types: the most cost-effective model, and the best-performing model. The system usually has a default option, possibly the most cost-effective, but we also leave the best-performing option for users to switch.
Because our application scenarios are quite many, such as TT ads, Paid Marketing, digital humans, or some narrative videos, educational videos, we run regression tests as much as possible ourselves, but you can't exhaust all scenarios. So the user data closed loop is very important to us. We actually have a team dedicated to monitoring whether these videos are actually used by users and placed across social media. This user-data-closed-loop feedback into post-training is very useful for optimizing our automatic model-switching routing.
Yi Shi: Let me answer about vertical-domain uncertainty control. In the vertical field of influencer marketing, having our AI game and negotiate with Influencers, the underlying general large model definitely can't directly make negotiation decisions well for me.
Essentially, we took all the core information our own operators communicated with influencers over the past four or five years — emails, WhatsApp, etc. — did a lot of vertical training, and built a small model (Small Specialized Model) that knows this industry best. So basically when doing concrete business negotiation and decisions, our private small model makes the judgments; the general large model does reasoning or general-text polishing on the periphery. So for us, the core of controlling uncertainty is doing well the private small model belonging to our own industry.
Core Discussion Three: How to Build an Agent's Memory System?
Host: Next question: if your work focuses relatively on Memory in multi-agent, can you share how to build this system?
Lux Li: As introduced, TiDB actually provides the whole Agent Infra; on memory we have a complete Solution. Recently at OpenAI's release event, a user also used our database for a memory solution. After observing with all users, we found Agent memory storage has big business-segment differences:
Some are work-type Agents; some products are general/general-entertainment Agents. The memory these two want to store is completely different. A General Agent may more want to record your personal Personal Profile, and Workspace Profile extending from it. A work-type Agent more purely wants to record work-related things, and may distinguish different Projects, even consider Memory Sharing across Projects. We have many production-grade users directly requiring memory to support sharing, and this knowledge and memory must pass down layer by layer like an org chart, because they think the knowledge junior employees consolidate is also useful for upper-level decisions. So memory isn't something a General just dumps in; it must be layered according to your own business.
Second is how memory is extracted and who extracts it. Especially for Agent platforms wanting self-evolution, each Agent is bound to a previous User. This User has their own work-preference tendencies or personal hobbies. This information is best recorded by the Agent itself. When cooperating with top Agent vendors, they tend to put the whole memory flow into the Agent Loop, letting the Agent itself do memory extraction and recall, and pull back useful memory at the Action stage.
For more vertical AI products, you also need to consider memory Privacy Isolation. When your users reach 100k, 200k, even millions, if multi-tenant isolation isn't done well it easily becomes a catastrophic privacy incident. We've put a lot of work into multi-tenant memory isolation. Memory determines the likes and prejudices your Agent shows; when building a memory system, you must be clear about what to store; a memory system highly related to business is a good system.
Jianfeng Lu: Our system also has quite a lot related to memory, such as enterprise business processes and policy Knowledge Base. But in building Agent systems, we found a difficulty is memory refresh and traceability.
Take a simple scenario: in March our company's promotion policy was such; by April, many March policies have expired. But some clients may look back and ask about that March deal. So throughout the process, besides putting time-sensitive data in the system, some expired or cold data needs MCP (Model Context Protocol) tool calls to trace back the historical info. There's a lot of detail handling here, and we also look forward to results from market players like EverMind in memory research.
Elliot Chen: This is indeed our comfort zone, because we've always focused on memory. Since release, our benchmark has been very high. Memory, in my view, has gone through several stages: the first stage everyone competed on how to effectively solve recall, saving users many tokens, very direct.
By this year, especially after self-evolving Agents like Hermes Agent rose, the playbook is completely different. We also make our Agent self-evolve; only when an Agent can self-evolve does it truly count as a living AI Assistant.
In our memory system, there are two very critical points:
Memory must be "private and human," which we call "born-with-it feel." What is born-with-it feel? When I say a sentence, the other party immediately understands what I mean. So our memory system does Preference layering on the user's past, then has a Forecast function to anticipate your future tendencies, making your private Agent more Proactive and private.
Enterprise-grade Knowledge Wiki and Reflection mechanism. The Wiki function can very effectively consolidate knowledge, to some extent replacing traditional RAG. At the same time, we have a trump-card feature called Dreaming (Reflection). The hot terminal large models all mention this now. It means in an Offline state, the Agent can use this feature to review and summarize the day's or previous stage's execution, thereby improving itself and achieving implicit self-evolution.
Benson Xu: The previous speakers were good; memory must be highly related to business. Let me share our Context and Memory storage practice in video. We actually divide into three layers of assets:
First layer is brand and product personalization data: including our clients' brand tone, product features, and a series of Enterprise-customized and highly personal assets.
Second layer is the data closed loop of ad materials and ad effect: the underlying data such as the CAC (customer acquisition cost) and Click of these materials after being placed, we all store. This later forms the user's own Agent Loop to automatically update and guide material generation.
Third layer is the user's own Workflow: we save the user's habitual workflow, generate their own unique Skill through Self-involve, so the user can call it repeatedly.
Though technically everyone calls this Memory, when talking with clients we often don't say Memory; we tell them: this is "user asset" exclusive to your enterprise.
Yi Shi: Let me talk about our memory implementation in the Creator Marketing field. AI communicating with influencers has a starting point: we need the brand or user to first pass us your website or some basic materials about your company. Our agent does intent recognition, understands who you are, then we can start the first round of communication.
In the whole Workflow afterward, AI communicates with influencers a lot: chat, price negotiation, even after reaching cooperation it stores some image and video materials. We comprehensively store the images, videos, and chat records produced throughout the workflow as brand assets.
In this storage process, we introduce another mechanism called the Auditing Agent. It cuts some unnecessary, valueless redundant content in the back-and-forth communication, keeping only the most core, valuable structured content in our memory base. The more the user uses it, the smarter its Agent gets; this is our practice in memory.
Endgame outlook: the multi-agent technical breakthrough you most look forward to this year?
Host: Thank you all; time is running out. Finally, please each use one word or one sentence to look forward to the Multi-Agent technology you most expect or think most needs to break through this year.
Jianfeng Lu: I think the most important is Reliability. Especially enterprise applications, multi-agent enterprise applications must be absolutely reliable. Whether the dialogue quality of answers, the output results, or in later version iteration, with back-end tools changed and models switched, whether the system remains as reliable after iteration. So the breakthrough I most look forward to this year is reliability.
Lux Li: Because we do infrastructure, actually those choosing TiDB include various top Agent frameworks, model vendors, and hardware vendors. The core metric by which they choose us is Stability. How to solve the complexity of Agents moving from Demo stage to production-ready? How to do privacy isolation well? How to do Scalability well? If you want to build a rigorous production-grade Agent, stability must be the most important.
Elliot Chen: My word here is achieving "Memory freedom." Whatever Agent or LLM you use, we hope to achieve cross-Session, cross-Agent memory continuity, so this Agent always belongs to you, with a true "born-with-it feel."
Benson Xu: We most hope to truly help clients directly Deliver a usable Final Video. This actually contains two keywords: one is the stability just mentioned, the other is Self-evolve. Marketing videos have a lot to do with timeliness, so its ability to build a Loop based on user data and continuously iterate and self-evolve is especially important.
Yi Shi: Based on our vertical going-global industry engineering practice, I think two things combined are especially important: first is Context Management, second is Workflow orchestration. Do these two well, and we can truly apply it to real work production for a certain vertical industry.