---
title: "Why Is Your Agent an “Artificial Idiot”? Six Industry Leaders Reveal the Painful Truth About Enterprise Deployment"
author: "Unique Research"
sourcePublication: "Unique Research Substack"
originalPublishedAt: "2025-12-02T09:16:48+00:00"
canonical: "https://ffcap.cn/en/research/src-20251202-01html"
source: "https://uniqueresearch.substack.com/p/src-20251202-01html"
language: "en"
---

# Why Is Your Agent an “Artificial Idiot”? Six Industry Leaders Reveal the Painful Truth About Enterprise Deployment

_Original · Unique Research · 2025-12-02_

_Historical-edition note: This preserves the original 2025 reporting, all four transcript sections, and the speakers’ contemporary roles, business figures, product comparisons and opinions; these are not current or independently audited performance claims. Company names without a source-provided English form are transliterated from the Chinese text, not asserted to be official English brands. The source’s illustrative probability argument assumes independent turns: 0.95 raised to 10 is approximately 59.87%, making the probability of at least one error approximately 40.13%. The transcript’s loose “a little above 50%” and “very likely” wording is retained as the speaker’s wording, not a different verified calculation. The annual-fee example does not specify currency. The opening infographic repeats the body’s accuracy, scenario-fit, tacit-knowledge, context, narrow-task and human-handoff points; it adds no separate dataset._

If you randomly stopped a few frontline employees at a company and asked, "Is that AI agent you've been using lately any good?"

Chances are, no one would answer immediately. Yet these employees have quietly given the Agents a nickname: "artificial idiots."

This is the reality inside many companies today: in PPT decks, agents are reinventing the business and defining the next-generation productivity paradigm; in actual roles, they become a pile of absurdly wrong answers, fabricated facts, and failures at critical moments.

People are being pushed to embrace agents while simultaneously inventing nicknames for them.

At a roundtable during the 2025 Beijing UniqueBloom, the participants simply put the unvarnished truth on the table. The moderator was Zhu He, CEO of Intelick, who leads AI investment, incubation, and international GTM at Yidian Tianxia. Across from him sat several frontline practitioners who work with agents every day: Zhai Xingji, founder and CEO of Yuhe Technology, which builds digital presales employees; Yang Jinsong, founder and CEO of Weilaishi Intelligence, who left Alibaba DAMO Academy to build AutoAgents.ai; Hu Xinran, founder and CEO of Yixuan Technology, who has helped Chinese manufacturers expand overseas since 2009; Zhang Wenhao, founder and CEO of Dudao Technology, which has worked deeply on AI agents for sales; and Sang Zhuohao, senior KA director at Focus Media, who deploys Agents from the enterprise buyer's perspective.

They put this practical question on the table:

Why do agents that look dazzling in a Demo become "artificial idiots" as soon as they enter an enterprise?

How far are we from genuinely trusting an agent?

What exactly is stupid about the "artificial idiots" described by frontline employees?

Zhai Xingji of Yuhe Technology went straight to the most painful point almost as soon as he began: accuracy.

They began building Agent digital employees in 2024, positioning them as presales solution specialists and serving leading customers ranked among the top companies in their industries. For a long time, they communicated just one sentence externally: an Agent's accuracy determines whether it is an employee or a joke.

In their view, an Agent is like a newly hired employee: if someone consistently scores only 60 points, you will certainly fire them; what you expect is stable performance above 90, with occasional mistakes still forgivable.

Many agent projects die according to the same script: the initial solution presentation is impressive and the boss approves it; frontline employees try it for two weeks, discover irrelevant answers and fabricated facts, and the project is ultimately shelved without fanfare.

It is not because no one loves innovation, but because no one wants to entrust their KPI to a digital colleague who scores 60.

More importantly, that score of 60 is often caused not by the model's capabilities, but by poorly understood requirements, inadequately decomposed logic, and badly designed architecture.

Zhai Xingji compared Agent projects with traditional information systems:

In the past, system projects investigated business processes and data;

Today, Agent projects must also take apart the mind of a business expert:

How does the expert make judgments? How are exceptions handled? Which parts are hard rules, and which are experiential intuition?

These things do not exist in any Excel file or PPT deck. They exist only in moments when an expert cannot quite explain something but simply knows it.

If this tacit knowledge is not uncovered, even the strongest model will ultimately be used as an "artificial idiot."

Yuhe therefore did something specific: rather than expecting to hire mature Agent engineers directly from the market, it built a systematic internal Agent-architecture training program that takes new hires from zero to independently designing Workflow, context-management, and Loop mechanisms. The models are already smart enough; what is lagging is the organization.

The Real Key Is Not an Insufficiently Powerful Model, but the Wrong Scenario

Yang Jinsong of Weilaishi Intelligence added a useful dimension to the phenomenon of "artificial idiots": the fit between time and scenario.

The team left Alibaba DAMO Academy to build AutoAgents.ai and adopted a global, full-stack approach from the outset: connecting the entire chain from underlying models, through the intermediate toolchain, to upper-layer business data. As a result, it works with major clients that impose exceptionally strict outcome requirements, including State Grid, China's three major oil companies, and major banks.

There is no room for experimentation in these settings. Deliveries must simultaneously satisfy security, permissions, controllability, and accountability.

Yang Jinsong's view is direct:

Many "artificial idiots" are not the result of bad models. They arise because you put a model into a scenario at which it was never good.

Take Text-to-SQL, the once intensely popular idea of querying data through natural language. In 2023, models were not yet sufficiently capable at code generation and semantic transformation, so the results were predictable. By the end of 2024, however, reasoning models had improved. With clear constraints and controllable table structures, the success rate of single-table Text-to-SQL had approached the threshold for practical use, making such scenarios worth revisiting.

In other words, the capability boundary of agents expands dynamically.

You cannot use a 2025 model to prove that a 2023 project was destined to fail, nor can you use a failure from 2023 to deny what may be possible in 2025.

Yang Jinsong's definition of an "artificial idiot" is, in fact, quintessentially engineering-oriented:

If you let a foundation model run blindly through a completely open task space, it will of course behave "idiotically." But if you give it a narrowed context and ask it to perform discriminative tasks or operate within clear boundaries, it can be highly reliable.

This is also a term he repeatedly emphasized: Tech-Market Fit.

Deploying agents today is no longer solely a matter of Product-Market Fit. The first question must be:

At the current level of technology, is this scenario truly suitable for an Agent?

If the answer is no, or if making it work would require an extremely large budget, waiting may be a more responsible choice than forcing deployment.

For Some Companies, the Scariest Problem Is Not Stupidity but Wildly Fluctuating Capability

Hu Xinran of Yixuan Technology approached the issue from a completely different angle.

His company is a long-established business founded in 2009, with 85 million monthly unique visitors across its network, helping Chinese manufacturing companies expand overseas.

It is an extremely labor-intensive business. The company has established more than 30 offices nationwide, and its employees traveled close to 2 million kilometers on customer-service visits last year for one purpose: genuinely helping customers build their businesses.

At a company like this, introducing agents is not about telling a new story. It is about transforming a business logic that has operated for more than 10 years.

Hu Xinran's fear is therefore entirely different from that of many AI-native companies:

We are not afraid that it may be a little stupid today. What scares us most is that it may suddenly become too smart.

If the model's capability is 1 today and a new release raises it to 10 tomorrow, that sounds positive. For a process-heavy company, however, it means every prior IT transformation, process optimization, and employee-training effort designed around capability level 1 may need to be rebuilt from scratch.

Conversely, an upgrade from GPT-4.5 to GPT-5 that was less impressive than expected was good news for them. Technological convergence meant they could safely make long-term investments at this level.

This reminds us that the agent economy operates on two timescales:

One is the model vendor's timescale, chasing one SOTA leaderboard result after another;

The other is the timescale of enterprise operations, focused on sustainable returns over 5 or 10 years.

On the latter timescale, a stable score of 8 can sometimes be more valuable than performance that swings between 9 and 12.

That is why many seemingly conservative companies are not actually opposed to AI. They simply have not yet seen a deployment approach that truly matches their own cadence and risk tolerance.

Multiple Dialogue Turns Mean Multiple Chances to Hit a Mine: The Brutal Mathematics of Sales Agents

Zhang Wenhao of Dudao Technology reduced the "artificial idiot" problem to a simple probability calculation.

The company began combining AI and marketing in 2012 and now focuses on agents for sales:

They must speak like real people on the phone, chat like real people on WeChat, and decide at the right moment when to hand the conversation to a human.

These may all sound like problems of human-like behavior, but the difficulty rises sharply once a conversation spans multiple turns.

Zhang Wenhao offered a striking calculation: suppose a foundation model answers a single turn with 95% accuracy. That already sounds excellent.

Yet even a simple sales conversation often continues for 10 turns or more.

What is the probability that not a single sentence goes wrong across those 10 turns?

Multiply 0.95 by itself 10 times, and the result falls below roughly 60%.

In other words, even with 95% single-turn accuracy, users are still very likely to encounter an "idiotic" answer over multiple turns.

Human memory is biased: the 9 correct sentences are taken for granted, while a single idiotic one diminishes the entire experience.

Faced with this reality, Dudao Technology did two things:

First, it continued raising single-turn accuracy, which is the foundation.

Second, it acknowledged that hallucinations cannot be eliminated completely and added a low-cost quality-control and correction process after every sentence.

Once a problem is detected, the output is either blocked or the conversation is transferred seamlessly to a human.

Its conclusion is:

The future of AI sales will necessarily combine people and machines.

What we replace is not the expert salesperson who can close complex enterprise accounts, but the large volume of repetitive, standardized outreach activities.

That may sound pragmatic, but it constitutes an important reality of the agent economy:

Not every step is suitable for an Agent; some are better assigned to more expensive human professionals.

The Enterprise Buyer's Sense of Reality: Do Not Imagine That One Agent Can Cure Every Problem

If the preceding speakers represented people who produce agents, Sang Zhuohao of Focus Media stood on the side of those who use them.

He offered two observations worth repeated consideration by anyone preparing to promote an Agent within an enterprise.

First, the most successful current deployments are not external-facing Agents, but internal knowledge Agents.

Consider meeting management: after a meeting, who turns scattered discussion, decisions, and action items into a project knowledge base that remains callable over time?

This work is highly structured and something no one wants to do indefinitely, making it well suited to an Agent.

Agents serving complex customers directly and handling extremely long contexts, by contrast, are currently less stable.

Second, excessively complex Agent architectures create performance surplus and information overload.

When Focus Media helped brands position themselves and write advertising slogans, it experimented with highly complex Deep Research processes for Agents:

It had them read large volumes of reports, competitor materials, and user research, aggregating the inputs layer by layer.

The result was that neither the human creative team nor the client accepted the advertising slogans produced after the complete complex process.

The problem was not that the model was too stupid, but that the context contained too much irrelevant noise. An incidental remark from the client could instead be interpreted by the model as a central requirement.

The team therefore stepped on the brake. It stopped imagining that one Agent could complete the entire journey from user insight to media strategy and creative output, and instead broke the task into small pieces so that the Agent performed only one task at each step.

On one occasion it only helps consolidate user research; next time it only organizes materials; after that, it only checks the logical chain.

Numerous people relay the work and perform checks between those stages.

The lesson is that current models are already sufficiently capable. The true challenge is whether you have the courage to admit that a model is suitable for only one small segment of the work.

Context Engineering: From Prompt Mysticism to Information Dieting

Over the past year, Prompt Engineering was the favorite topic whenever people discussed agents.

Yet this conversation suggested that a new term has quietly taken the baton: Context Engineering.

Sang Zhuohao offered a vivid analogy: you encounter Focus Media CEO Jiang Nanchun in an elevator and have 60 seconds to explain your brand's background, users, and needs before asking him to produce an advertising slogan on the spot.

Most people would hesitate to use that slogan, because 60 seconds of context cannot possibly contain the problem's complexity.

Many people assume that if a Prompt is long and arcane enough, an Agent can understand all the background.

The reality is that the context you provide should not simply be larger; it should be more relevant.

Focus Media's practical path was roughly as follows:

In the first phase, it tried giving the Agent every potentially relevant piece of information and allowing it to decide what was missing and how to supplement it.

The team discovered that the Agent often focused on the wrong points and treated nonessential information as the main thread.

In the second phase, it added a refinement layer above the raw materials.

One layer of logic first extracted the key information, and that information then became the context primarily called by the Agent.

As a result, the team spent 90% of its time on context engineering and only 10% on the Prompt.

Weilaishi Intelligence practices another form of context engineering. It divides context into several functional modules: memory-based user preferences and history, data returned by tool calls, and background knowledge obtained through retrieval augmentation.

Static elements use Prefix Caching, while dynamic elements are continually updated as the task advances.

The key is that the team has designed evaluation methods specifically for context engineering, allowing it to quantify which forms of context partitioning are more effective.

Yixuan Technology, meanwhile, offered a cost-sensitive perspective: under its annual-subscription model, every additional foundation-model call is paid out of its own pocket.

It therefore asks the reverse question: what does not require a foundation model at all?

Anything already available in the knowledge base and retrievable semantically or structurally goes through the cache first.

Only the 10%-20% of work that genuinely requires reasoning is left to the foundation model.

If prompt engineering resembles writing an instruction manual for a foundation model, context engineering resembles putting it on an information diet with balanced nutrition: removing excess fat and retaining only what truly supports the decision.

Seeing Accurately, Retrieving Accurately, and Executing Accurately: A Systematic Method for Reducing Hallucinations

Returning to the issue that causes everyone the most anxiety: how can hallucinations be controlled?

Yuhe Technology's summary was concise yet substantive: accuracy.

First, see accurately.

If the input is inaccurate, flashy output is useless. In enterprise settings, seeing a problem clearly is not as simple as running OCR on a document;

the logical relationships in flowcharts, layout information in PPT files, and relationships among tables must also be reconstructed. Otherwise, the Agent sees only fragmented information and naturally makes things up.

Second, retrieve accurately.

Many people assume the difficult part of RAG is the vector database, but the true challenge is how the question itself should be defined.

Yuhe divides this into Agent Planning and Agent RAG: first make the expert's thought process explicit and create a general search strategy, then have the Agent locate and combine knowledge according to that strategy.

When a new client arrives, its materials can simply be provided, and the Agent retrieves information according to expert logic rather than rummaging through it based on model intuition.

Third, execute accurately.

Even if you use the newest foundation model and directly assign it a sales task, it will still behave idiotically.

Excellent salespeople possess large amounts of tacit knowledge: when to ask another question, when to conclude, and when to remain silent.

This knowledge often exists in no document. It resides only in situational memories that are difficult to explain but in which something simply must be done a certain way.

Yuhe makes this tacit knowledge explicit through extensive interviews and co-creation, then encodes it in a Workflow,

uses synthetic data to fine-tune the model, and turns the result into a genuinely reusable digital employee for the role.

Together, these three elements form a complete anti-hallucination system:

They show that hallucinations are not solved through one extremely long prompt, but reduced through an entire system of engineering and knowledge methodology.

Crossing the Trust Singularity: From "Artificial Idiot" to Trusted Colleague

After all this discussion, we return to the phrase from the beginning: the trust singularity. It does not mean a day when a model suddenly becomes omniscient and omnipotent,

but rather a turning point inside an enterprise when most employees become willing to entrust agents with some real responsibility.

Before this point, an agent is a toy used for Demos, promotional copy, and the boss's presentation PPT.

After this point, it truly becomes part of the organization:

It answers calls, responds to customers, writes proposals, manages the knowledge base, and even participates in the KPI system.

This conversation suggests several possible paths toward the trust singularity:

First, design the Agent as a person rather than use it as a plugin.

It needs clear job responsibilities, assessment standards, a training system, and an advancement path,

rather than being a universal robot burdened with every kind of task.

Second, acknowledge that human-machine collaboration is a long-term state, not an effort to replace people completely with AI.

Where a human handoff is needed, design a seamless transition mechanism,

so that users do not feel, at the most critical moment, that they are being treated cheaply.

Third, restrain ambition when selecting scenarios and concentrate on going deep in one or two small scenarios with sufficiently high value.

Even if it is only a meeting-notes Agent, a digital presales employee, or a website-configuration assistant,

if it genuinely saves the business time and generates profit, it is more persuasive than 10 flashy concepts.

Fourth, accept the pace of technological evolution and give the organization a stable learning period.

Not every model upgrade justifies overturning existing processes.

In many cases, reliably deploying the previous generation's capabilities is more economical than chasing the newest generation without being able to implement it.

One final point:

The nickname "artificial idiot" carries the emotional release of frontline employees. It is not merely ridicule of technology, but resistance to the feeling of being put through unnecessary disruption.

To truly enter the era of the agent economy, we need not only stronger models, but also a patient understanding of people, organizations, and the cost of trust.

When employees one day cease treating an Agent as a new toy imposed by the boss,

and instead treat it like a reliable new colleague—willing to teach it, supervise it, and let it work independently—

only then will the so-called trust singularity truly have been crossed.

More Details from the Conversation

Part One: Opening and Introductions

Zhu He: Let me introduce myself first. I currently oversee AI investment, incubation, and mergers and acquisitions at Yidian Tianxia, as well as GTM (Go-to-Market) programs that help various companies in the AI industry expand overseas. I am essentially the chief expert in this area.

Today's discussion begins with our first keyword: "artificial idiot." Everyone is discussing agents, so that is our first topic. In a Demo, today's AI Agent always looks relatively perfect, whether the result was cherry-picked or the Demo was deliberately faked. Yet during real adoption, it breaks down and encounters all kinds of problems, whether deployed in a small team or rolled out broadly across enterprise departments.

We also have an AI coach from Chaos here today. In our conversations offstage, we found that while coaching many companies, frontline employees often rated agents very poorly. We therefore want first to unpack the phrase "artificial idiot": does an agent behave idiotically because current model capabilities remain insufficient—for example, because GPT-5 is still inadequate and we must wait for GPT-6 or GPT-7—or because the Agent architecture is wrong during implementation, or because the engineers lack sufficient technical capability? What exactly makes people call it an "artificial idiot" rather than an intelligent agent?

Let us begin with Mr. Zhai and proceed in order.

Part Two: Roundtable Discussion—Why Do "Artificial Idiots" Appear?

Zhai Xingji: All right. Thank you, Mr. Zhu, and thank you, everyone. Let me introduce myself. I am Zhai Xingji, founder and CEO of Yuhe Technology. We are an AI Native company founded in 2023, and we began developing Agent digital employees in 2024. Our progress through this exploration has been relatively smooth. Today, we primarily build Agent digital employees positioned as presales solution specialists. Put simply, this Agent uses technical expertise to develop solutions, architectures, and designs that help enterprises sell things to top-tier customers. Our core customers are also leading companies ranked among the top in China or worldwide in their respective sectors.

On the "artificial idiot" issue, I think the revised topic is excellent and highly relevant to us. When we first began building To B Agent digital employees last year, we identified the greatest pain point: Agent accuracy.

For more than a year, we have consistently communicated one core value externally: Agent accuracy matters. At last year's WAIC World Artificial Intelligence Conference, we had no product or Demo, only an early solution and a technical team. We added many people on WeChat at the event and had excellent conversations for one central reason: we asked only one thing. Have you deployed an Agent inside an enterprise? If so, was insufficient accuracy the greatest pain point? Did the business abandon it as soon as people used it? Did the initial presentation feel excellent, only for the boss to conclude that the result was poor and become unwilling to promote new projects?

That point strongly resonated with people and became the foundation of our Agent work. From last year through the middle of this year, much of our work therefore focused on raising end-to-end Agent accuracy high enough.

We hold a basic belief: an Agent digital employee is like a human employee you recruit. If you hire someone who performs at only 60 points, that person will be fired; you must require 90-point performance. Reaching 90 involves many difficulties, but I believe the greatest challenges lie not in the model, but in requirements analysis and design.

Requirements analysis: Agent projects differ greatly from traditional digital projects. Traditional digitalization investigates business processes and data. An Agent project must additionally engage human business experts to uncover their thinking logic, Know-how, SOP, and response patterns across different scenarios, then try to turn these into data and models.

Architecture design: Even after understanding the requirements, you must be capable of designing the Agent architecture. Our early delivery projects also encountered problems and had insufficient accuracy. The fundamental issue remained the team's limited ability to design Agent architectures—very few people can build an Agent well, understand the capability boundaries of different models, and properly manage context, connect Workflow, and design Loop mechanisms.

We therefore built a complete internal training system for Agent-architecture capabilities. We do not expect to hire experienced people from the market, because we found that people claiming 2 years of experience had very limited capabilities, while our own excellent employees could quickly acquire the needed skills after 2 or 3 months of internal training.

Zhu He: Your conclusion, then, is that model capability is temporarily sufficient. More importantly, companies must understand the requirements, design the architecture properly, and establish a system for training employees. Thank you. Let us hear from our second speaker, Mr. Yang.

Yang Jinsong: All right. Let me first introduce myself. My name is Yang Jinsong, and I am the founder of Weilaishi Intelligence. Our project's English name is AutoAgents.ai, and it was founded in June 2023. As the name suggests, from our first day we have provided integrated products and services to the global market. Our background is a team of 3 co-founders from Alibaba DAMO Academy's Tongyi Lab, and our path is relatively full-stack technologically: from underlying models to the intermediate Agent toolchain and then upper-layer business and data, we seek excellent integration to improve final Agent delivery outcomes.

Returning to the "artificial idiot" issue, I believe the central question is whether the demand scenario is something at which agents excel.

When foundation models first appeared in 2023, many companies idealistically wanted to develop Text-to-SQL, converting natural language into SQL to query and analyze data. At that time, foundation models were relatively weak at code generation and semantic transformation, so their agents performed poorly. By the end of last year, improvements in reasoning models and capabilities had gradually unlocked new scenarios.

Why do agents become "artificial idiots" in some scenarios? Because our expectations are too high and we expect them to do everything. The core of successful deployment is the fit between scenario and technology: Tech-Market Fit.

The clients we serve today are highly serious major organizations such as State Grid, China's 3 major oil companies, and large banks. They impose the strictest delivery requirements, whether for security or permission controls. The applications nevertheless currently perform well because we first determine whether the desired business scenario can be satisfied effectively at the present stage of technology.

For example, if you ask a foundation model to perform a completely open-ended task, it will certainly behave idiotically. But if you give it context and ask it to perform a discriminative task—for example, providing a rule and judgment requirements and asking it to analyze accordingly—it can perform extremely well. Similarly, code-generation capability has now improved, and the success rate of single-table Text-to-SQL exceeds 99.9%, so the scenario is viable this year.

Therefore, turning an agent into a usable, highly accurate system depends fundamentally on matching it with the technology. If the two do not match in the short term, results in a particular domain can still be improved through methods such as constructing a dedicated environment and reinforcement learning, though this depends on the client's budget.

Zhu He: Understood. Your answer has given me considerable inspiration. In effect, you introduced a dynamic time dimension: an Agent's capability boundary continually expands and must match the current stage, such as GPT-5's capabilities. Second, if capability is insufficient, the upstream process may be inadequate and the Context may not have been prepared fully. I am now worried about the next 3 guests: the horizontal and vertical time dimensions have both been covered, so what can they discuss? Let us welcome our third speaker, Mr. Hu.

Hu Xinran: I am Hu Xinran of Yixuan Technology. Compared with many AI companies, ours is relatively old, having been founded around 2009. We do something fairly traditional: help Chinese manufacturing companies expand overseas. Our network has approximately 85 million monthly unique visitors.

We are the internet's manual laborers. What do I mean? We established roughly 30 offices nationwide and visit customers one by one to discuss business. It is exceptionally hard work. Last year, we calculated that employees making on-site after-sales service visits traveled close to 2 million kilometers.

The preceding 2 guests represent companies born because of AI and operating within AI, and they need to explore the fields in which AI has limitless potential. Our situation is more definite: we use AI to transform our own business and upgrade our commercial logic.

Our perspective on this topic is therefore different. We do not fear stupidity; what frightens us most is instability.

What I fear most is that the foundation model has capability level 1 today and becomes 10 after a new release. Then every IT investment and every process transformation I made today becomes worthless. I do not care whether it knows things or is intelligent. I care whether it is stable. If its intelligence cannot reach a stable state, I am very afraid.

Consider GPT-4.5 to GPT-5: the change was not as large as everyone expected. That creates considerable pressure for exploratory companies, but it is acceptable for a company like ours. It means the technology has stabilized, so perhaps I can promptly undertake many changes that optimize the company's business logic. Instability is what frightens me most.

Zhu He: Our guests are growing increasingly formidable and have begun challenging the question's pivot point. We very much look forward to the final 2 guests. Our fourth speaker is Mr. Zhang.

Zhang Wenhao: Hello, everyone. My name is Zhang Wenhao, and I am from Dudao Technology. Mr. Hu mentioned that Yixuan Technology is a long-established company; in fact, we were also founded early, in 2012. At the time, I was pursuing a doctorate at Tsinghua University and mainly conducting AI research, continuously studying how AI could be combined with marketing. Dudao Technology now primarily builds agents for sales.

In sales, we focus particularly on several areas:

Human-like behavior: Voice or telephone communication increasingly resembles a human. TTS has recently improved substantially. In China, sales mainly depends on telephone and WeChat, and we focus on behaving more like a person on WeChat while communicating with users automatically.

Interface-calling capability: The central issue is when to transfer to a human. This requires accurately predicting whether the current stage calls for an external tool or a handoff, and we use some reinforcement-learning technology in this area.

On the "artificial idiot" issue, I think the preceding guests made excellent points. Selecting the right scenario is the most important part of deployment.

Dudao Technology views sales as a long chain of multi-turn conversations. Even simple sales requires at least 10 dialogue turns. If a foundation model's accuracy on a single call is 95%, multiply the probabilities: 0.95 to the 10th power is only a little above 50%. This means that after 10 dialogue turns, an idiotic response remains very likely.

Our core considerations are therefore:

Improve single-response accuracy: As Mr. Zhai of Yuhe noted, optimize the underlying system to solve accuracy problems in the model itself.

Quality inspection and correction: Foundation models cannot completely avoid hallucinations; the mechanism makes that inevitable. Since hallucinations will occur, how can we inspect each output sentence with a low-cost solution? If a problem appears, block the result or modify it directly.

I believe AI sales combines people and machines. AI will struggle to replace a highly mature salesperson handling major offline accounts, but it is well suited to highly repetitive scenarios such as insurance telemarketing and WeChat sales.

This also invites comparison with AI programming. Cursor existed before GPT-4 and performed poorly at the time, but because its productization was strong, it seized the opportunity when Claude released a new version and the underlying capabilities improved. The lesson for us is to examine TMF (Technology-Market Fit). If we predict that a technology is immature today but will mature within the next 6 months or 1 year, the scenario remains worth pursuing. When the underlying foundation model becomes good enough, our market opportunity will arrive.

Zhu He: Understood. Mr. Zhang supplemented the discussion of application scenarios and added a quality-inspection and optimization layer on the back end. The pressure now falls on the final speaker, Mr. Sang.

Sang Zhuohao: Let me first introduce myself. My name is Sang Zhuohao, and I am from Focus Media. I am responsible for KA sales and also lead Focus Media's AI efforts, so my background differs from those of the preceding speakers: we are an enterprise buyer that uses Agents.

During use, we found that because the customer situations we face are exceptionally complex and differ greatly from traditional telemarketing, our context length far exceeds that of ordinary telephone sales.

Our most successful internal deployment is not an Agent that directly serves customers, but a meeting-management Agent. After a meeting, how do we turn the knowledge into a project knowledge base that can support future negotiations and proposal development over the long term? It performs relatively well as a knowledge assistant.

We have also developed many Agents that help brands with positioning and advertising-slogan creation. We found that when an Agent architecture becomes exceptionally complex, it instead creates performance surplus.

Performance surplus means that after an Agent completes N rounds of Deep Research and aggregates the context, neither we nor the client accept the resulting advertising slogan. The reason is that too much AI participation in the process generates far more contextual noise than we can control.

We have therefore simplified our agents, asking them to assist only with the small, immediate task in each meeting or proposal round. We do not expect them to complete everything in a single pass from user insight through media analysis and competitor analysis—that produces disastrous results, and people cannot align with the system.

Our current approach is this: when conducting user research, we put everything into the user-research Agent for processing, extract the result, and then proceed to the next round. Many stages in between involve substantial human participation. In our enterprise setting, model capability is already more than adequate; the challenge is instead how to simplify the context.

Part Three: Context Engineering vs Prompt Engineering

Zhu He: Understood. That is very specific, and it introduces our next topic. You just mentioned that context can be excessive or too long. Since the beginning of this year, the term Context Engineering has become popular. Some institutions, such as Anthropic/Menlo, have also published videos or papers discussing context offloading. In the past, we spoke more often about Prompt Engineering.

Which approach do you use more in your scenarios, or do you combine them? Please share specific examples.

Sang Zhuohao: Over the past 3 months, we have actually been addressing context-engineering problems. At first, people easily understand an Agent as a Prompt combined with a Chatbot system.

Let me give an example. You may know our CEO, Jiang Nanchun, also known as Mr. Jiang. Suppose you encounter Mr. Jiang in an elevator today and have 1 minute to introduce a brand, hoping he will create a powerful advertising slogan for you. Would you dare use the slogan he produces?

Having surveyed many companies, we found that most said they would not. Why? Because 60 seconds is insufficient to explain everything and insufficient for this superbrain to understand my context.

We therefore studied exactly which contexts were needed for an output to understand the background fully.

At the time, we created a List and asked AI to automatically analyze which information gaps existed and fill them, placing all information in a file pool for it to call. The result was still poor, because the information it called did not align with our intended direction. A client might casually mention a nonessential point, yet the model treated it as the focus.

In phase 2.0, we therefore added a refinement layer above the foundational corpus and used information from that layer as the key material to retrieve. We may spend 90% of our time on Context Engineering and 10% on the Prompt.

Zhang Wenhao: AI sales also imposes high context requirements, because each conversation leaves user behavior, including purchasing traits, emotions, and history.

We need to solve 2 problems:

Extract effective data: Most dialogue data, including small talk and random conversations, is invalid. We must accurately extract key features from massive volumes of data.

Data compression and prediction: As data grows, how do we compress it and predict the next sales action based on the current state?

We write some data into the Prompt and place some in RAG. When invocation costs are high, however, we use reinforcement learning based on vertical scenarios to train scenario data into the model and reduce invocation costs.

Zhu He: Understood. Let me ask a direct follow-up: put simply, how do we solve hallucination problems?

Zhang Wenhao: Hallucinations cannot be completely avoided. They arise because the model has a low density of content from vertical domains. The solutions are:

At the data level: obtain high-quality data covering every Corner Case, just as autonomous-driving systems are trained.

Data synthesis: use language generation to simulate additional scenario data for model training.

Human handoff: Even these measures cannot guarantee 100%, so the key is finding the right time for a seamless transfer to a person.

Hu Xinran: We probably use prompts more than context because we address established businesses.

For example, in AI website creation, a user enters, "I want a highly technological-looking website." That is only the user's prompt. You must translate it into a System Prompt that AI can understand: what is the industry? What is the product? What are the color parameters?

Internal professionals must tune this. We tried opening it to customers and found that customers could overwhelm the AI. We therefore operate more by centering the user and translating the user's vague language into AI's professional language.

Zhu He: Understood. Scenarios often involve considerable repeated thinking. Ant Group recently launched a product that offers refunds for Tokens used to regenerate because of hallucinations. What is your view on those Tokens used for regeneration and rethinking? Should you build a caching mechanism or charge customers for them?

Hu Xinran: That depends on the business model. If I charge a markup per Token, of course I want you to use as many as possible. But we sell annual services, for example at 50,000 per year (currency unspecified in the source), so every cent I save becomes profit.

We therefore must use caching mechanisms, including semantic caches and knowledge-base caches. If something is already in the knowledge base, why ask a foundation model to do it? We can limit the work performed by foundation models to roughly 10%-20%; otherwise, we pay all the costs ourselves.

Yang Jinsong: We use more dynamic context engineering.

The Prompt is relatively fixed, while context engineering addresses Long-tail scenarios. We need to load several blocks of content dynamically:

Memory: user preferences and historical records.

Tools: results returned by APIs for search, weather checks, flight searches, and similar tasks.

Background knowledge: IDs or documents added through RAG for a particular step of a task.

We divide context into several blocks: the fixed part uses caching through Prefix Caching, while the dynamic part is generated as tasks are executed. This requires an evaluation system to determine whether context engineering is effective.

Zhai Xingji: From our perspective, controlling hallucinations comes down to 3 things: see accurately, retrieve accurately, and execute accurately.

See accurately (Input): Driving while nearsighted and without glasses is dangerous. We must ensure that the Agent's input information is sufficiently accurate. This requires extensive upfront data processing to reconstruct complex documents, flowcharts, tables, and layout information in PPT files, rather than merely OCR text.

Retrieve accurately (Retrieval):

Agent Planning: Defining a question is harder than finding an answer. An Agent must be capable of planning.

Agent RAG: Experts in a role have their own methodology. We preconstruct strategies for categorizing, processing, and retrieving common knowledge. When a new client arrives, it only needs to provide the materials, and the Agent can find information according to expert logic.

Execute accurately (Execution): Even the newest foundation models such as Gemini/Llama behave idiotically when directly asked to work as a Sales Agent. A model is a product of data, while an expert's tacit knowledge, or Implicit Knowledge, does not exist in explicit documents.

Through research and communication, we make the tacit knowledge in experts' minds explicit, turn it into a Workflow and synthetic data, and then perform Fine-tuning work. Only this complete systematic method can potentially reduce hallucinations.

Part Four: Investment Game (Ending)

Zhu He: Excellent. The three forms of accuracy are memorable. Let us end with an entertainment round: if each of you were an investor, which of the other guests would you invest in, excluding your own company? Mr. Sang is excluded because he represents an enterprise buyer.

Zhai Xingji: That question is too likely to make enemies... I would invest in Qingsong, that is, Yang Jinsong, because we both come from technical backgrounds and hold similar views.

Yang Jinsong: I would feel a little awkward investing in Xingji, that is, Mr. Zhai... I might invest in Mr. Sang at Focus Media.

Because an Agent must be integrated with an industry scenario, and Focus Media has accumulated extensive data, which is a strategic barrier. A technological lead lasts only about half a year; barriers such as data and reputation are what matter.

Zhu He: He is an enterprise buyer. You cannot invest in him.

Yang Jinsong: Then I would like to invest half in Xingji, that is, Mr. Zhai.

Zhu He: Understood. Regardless of whom you invest in, you believe scenarios and data matter more. Mr. Hu?

Hu Xinran: If I had to choose, I would invest in everyone, or perhaps no one.

Today's AI companies are all building at the application layer. The height of the barrier determines whether they have investment value. Within what I can currently see, very few AI applications possess genuine barriers. They may constitute businesses, but do not necessarily have investment value.

Zhang Wenhao: I would probably invest in Mr. Zhai of Yuhe Technology.

We are optimistic about the broad direction of AI sales. Apart from AI programming, AI sales is the easiest field to deploy and has enormous market potential. We share the same broad direction but serve different customer groups, which would hedge risk.

Sang Zhuohao: I am in a rather detached position now. I would prefer to speak more with Mr. Zhai and Mr. Zhang and invest half in each. I still want to solve problems in vertical sales scenarios, including how to win KA enterprise accounts.

Zhu He: All right, the music is urging me to stop. Many questions remain unasked. I hope we will have another opportunity to exchange views. Thank you, everyone!

---

Original publication: https://uniqueresearch.substack.com/p/src-20251202-01html
On-site reading page: https://ffcap.cn/en/research/src-20251202-01html
