
Original · Unique Research · 2026-08-10
Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening essay, six insight sections, speaker list, and the full roundtable transcript. The title's "Video Agent Is Dead" is attributed to Liang Wei as a challenge to specific shell-and-style generic Agents, not a consensus view of all five panelists. User, revenue, and technical figures are speaker self-reports, not independently verified findings.
Original · Unique Research · 2026-08-10 18:29 Shanghai
Is the AI Video Agent Dead?
AI Industry Observation
"What died isn't the Agent; it's the generic Agent that puts on a shell and goes to market."
The person who said this isn't an AI hater — it's Liang Wei (梁巍), co-founder of MovieFlow. His product has 1.5 million users across 174 countries; the pro version, launched this past May, has served over 100 Hollywood and Chinese film and TV crews. Someone who makes a living from AI video, publicly announcing video Agent death.
And it didn't stop there. On the same panel, Luo Yuan (罗媛, known as Shanzhu/山竹), AI video super-creator and Executive Director of Huanjing Jiyuan (幻境纪元), added a knife: "I don't let my employees use Agents; it increases communication costs."
This was an AI Agent roundtable at the Unique Conference, hosted by Jiang Zhiqiang (江志强) of Heli Capital (和利资本), with five people on stage who are all using AI to run content businesses. By rights, this should be the occasion most likely to sing Agent praises. Instead, five people took turns dismantling the year's hottest tech buzzword.
But here's what's interesting: for all the dismantling, these five are all quietly making money with Agents.
What are they really playing at? I listened to the whole thing and found the answer far more interesting than the question "are Agents useful or not."
"Should We Distill Wong Kar-wai?"
The direct trigger for Liang Wei's opening salvo was another wave of WeChat articles pushing AI video Agents and various director-style skills. His original words were blunt: this is just like last year when GPT came out and everyone had Ghibli style — it burned for a few days and was over.
"Did they get the director's copyright? On what basis do they use Wong Kar-wai's name? Should we distill Wong Kar-wai? What you distill is just a shell — you distill the visual feel of Wong Kar-wai's movies, but what does that have to do with storytelling? There's no necessary connection between a good tool and a good work. Give you an IMAX camera, you're still not Nolan."
Shanzhu's angle was different — the gripe of a pure practitioner. As a super-creator who has touched nearly every tool on the market, she tried using Agents to make video and concluded it's worse than doing it herself. Communicating with AI is laborious; most of the time it's like communicating with an employee who may not understand you, increasing communication costs. Plus, conversational content generation causes secondary iteration images to break up and get noisy, requiring rework. So she simply drew a hard line: employees are banned from using them.
One building a platform, one making content — reaching the same conclusion from two directions: the current batch of generic video Agents is a liability for people who actually do the work.
"Right Now, Only One Type of Agent Works"
Wang Bihao (王碧豪), co-founder of Xin Ying Sui Xing (心影随形), explained it thoroughly. His company serves over ten million gamers globally with AI companions and AIGC game communities.
His logic chain is clear. Why is it called an Agent? Because it has autonomous decision-making capability. The key to autonomous decision-making is the loop: receive a task, decompose it to various professional skills, synthesize, and finally have a reviewer judge how well it was done, feeding feedback back in.
The problem is this reviewer.
"In my view, right now there's only one type of Agent that works: the coding Agent. All other Agents actually don't hold up. Why? Because there's no reviewer — no one can judge whether this Agent did a good job or not. Coding works because the machine can clearly tell you whether the code is correct and runs. Whether a film is good, whether a video is well-made — only the audience can tell you, and it's subjective."
A friend of his has an even more pessimistic view: all Agents are losing their user base, because users fall into three categories. First, people with no ideas — give them an Agent and they still don't know what to do. Second, people with ideas but who can't write prompts well. Third, people who have ideas AND can write — they just go straight to Claude Code and Codex, why use your shell product? By this logic, Agent entrepreneurship is basically annihilated.
But Wang Bihao turned the corner: within the annihilation argument lies the next opportunity. People with no ideas — can they give feedback through light interactions like clicks, swipes, and likes? People who can't write prompts well — can someone bring the magical experience of Claude Code to where ordinary consumers can reach it? As for why third-category experts stay in vertical products, his answer is: building loops in different domains contains business-integration secrets that require true domain experts to uncover that domain Agent's truth. Note this last sentence — it's the key to the entire roundtable.
The Person Who Dismantled Most Hard Is the Most Addicted Himself
Next came the most contrastive moment of the whole panel.
Wang Bihao, who just said "only one type of Agent works," then laid out his own usage: he's a heavy Claude Code and Codex user, averaging 300 to 500 commits to his projects per day. In his own product, the play is even more aggressive: after users generate videos, games, and text, an Agent reviews it, the user also reviews it, likes and retention feed real-time into a judgment system, and if it doesn't pass, it gets sent back for rework. The Agent even rewrites code itself. Essentially, he doesn't disbelieve Agents — he disbelieves Agents without a referee.
Yin Tianming (殷天明), founder of Longyuanji AI (龙渊纪AI), takes another path: distilling people into Agents. He's former product VP at 37 Interactive Entertainment, 20 years in gaming; one co-founder managed ad buying for over a decade, handling tens of billions in spend; another spent over a decade at Tencent in publishing. Their approach is spoon-feeding decades of industry experience to AI. Dramas split into male-oriented and female-oriented, cultivation and transmigration; games split into MMO, card, SLG — each genre gets different playbooks, otherwise it's just chaos.
"What died is the generic Agent that puts on a shell and pinches a director style to sell; what lives is the digital employee that has industry know-how distilled in, clear referee standards, and accountability for ROI. Both are called Agents — one exists in WeChat articles, the other in financial numbers."
Why Is AI Video Only 15 Seconds?
The highest information-density three minutes of the panel came from Wang Bihao's technical breakdown. The question starts from a pain point every creator has: why can today's AI video only generate 15-second clips? Want 4K and you must cut to 4–5 seconds.
His explanation: this generation of models is all based on DiT architecture — Diffusion for frames, Transformer for time series. Parameter and compute requirements are gold-vanquishing-beast level; mainstream models have 20B to 30B parameters, hitting the architectural ceiling. To break through consistency, you need to train a world model that understands long shots, requiring 1 billion videos. And reality is, data is extremely scarce.
"Why does generated AI video — when someone opens a door, walks in, and walks out — come out with a different person, different clothes, different environment? Because in the training corpus, when a film or TV show opens a door, the perspective and lighting all change, so AI thinks: oh, this is how the world is."
The room went quiet. Others were talking about Agent concepts; he was talking about the physical constraints of training corpora. His solution is clever: game worlds are full of long takes — first-person perspective runs an hour continuously — so Xin Ying Sui Xing is using game footage to build long-shot corpora. That's why he dares make the "only one type of Agent works" judgment — the person saying it is actually running the numbers.
A Cat That Holds a Grudge
There was also a wild card at the table: Nie Yonggang Sawyer (聂勇刚 Sawyer), product lead at Meowster. They make an AI cat, targeting the Japanese market, with nearly 100,000 global users, 93% in Japan.
Why not human-shaped companionship, why a cat? His reasoning: cats have uncertainty — you can't guess what they'll do next; they might surprise you or do something annoying. How annoying? "If you ignore it for a week or two, there's a 70–80% chance it won't want to interact with you either, because it feels you can't give it the care it needs." All the cat's feedback is generated in real time by AI, and it has a good memory — it knows you're busy writing in the evening and busy working in the morning.
"If Jarvis is always at your beck and call and perfectly obedient, it's just a fancy remote control. A cat that holds a grudge and throws a tantrum feels more alive. The emotional companionship track may not be selling intelligence but uncertainty."
Token Consumption Is the Real Commercialization Litmus Test
Host Jiang Zhiqiang offered a metric: the industries represented here — short dramas, gaming, AI entertainment, companionship — are all high-stickiness, high-token-consumption scenarios. Concepts won't settle anything; token consumption is the litmus test for commercial viability.
Liang Wei left everyone with a closing thought: the tool layer is already very simple — anyone can hand-build a workflow in a day or two. The hard part is whether the workflow has accumulated creative assets and real know-how.
Back to those two opening provocations. Shanzhu doesn't let employees use Agents; Yin Tianming distills three co-founders into digital employees. Both are called "employees" — one is banned, the other treasured.
"The difference is only one thing: do you have something worth teaching it?"
More Conversation Details
Speakers
Liang Wei, Co-founder, MovieFlow
Wang Bihao, Co-founder, Xin Ying Sui Xing
Luo Yuan (Shanzhu), Executive Director & AIGC Multi-Platform Super Creator, Huanjing Jiyuan
Yin Tianming, Founder, Longyuanji AI (龙渊纪AI)
Nie Yonggang Sawyer, Product Lead, Meowster
Host
Jiang Zhiqiang, Venture Partner, Heli Capital AI Fund (和利资本)
Jiang Zhiqiang: Hello everyone, I'm Jiang Zhiqiang. I've been in this generative AI wave since it started in 2022 — early image generation, video generation, multimodal, to multi-agent architecture applications. Before becoming Venture Partner at Heli Capital, I did AI consulting and development in enterprise services for over two years. The organizers asked me to host this topic because over the past six months I've been jointly incubating AI short-drama self-production with a listed film company, working on premium content and Agents in this space. Let me ask each guest to introduce themselves. Starting with someone I already know fairly well, Liang Zong.
Liang Wei: Hello, I'm Liang Wei from MovieFlow. We're an AI video workstation for global creators, and this year we launched a pro version. Previously mainly serving global C-end users; this year's pro version is the MovieFlow Studio, launched in May. Studio currently serves major Hollywood and Chinese professional film and TV companies, professional creators, content companies, and enterprises. Because the workstation provides a lot of professional, detailed, creation-oriented capabilities — not the C-end one-click generation. In short, on the AI video layer we're an application-layer delivery platform, currently focused on truly delivering AI video generation capability and film/TV industry FDE services.
Wang Bihao: Hello, I'm Wang Bihao from Xin Ying Sui Xing. This is my second time on the Unique stage; last time was in Beijing. Our company serves gamers globally — through Doudou Game Companion and Hakko AI we've gathered over ten million global gamers, providing AI companion services based on visual recognition and real-time game-screen recognition with voice companionship, and building game communities. We're also expanding more multimodal services, including using AI to generate games, comics, and short dramas, expanding AIGC services within the gamer community. We have a concept called "consumption is generation" — users generating games and videos is itself the consumption of this AI entertainment experience, which we keep practicing. Glad to share some thoughts today.
Jiang Zhiqiang: Thanks, Wang Zong. Next, Yin Zong.
Yin Tianming: Hello, I'm Yin Tianming, founder of Longyuanji AI. Like Wang Zong, this is my second time at a Unique event. Our company is building a next-generation content creation and intelligent ad-delivery engine driven by ROI. The essential goal is to distill our nearly 20 years of ad-buying experience into an executable AI system accountable for ROI. We bundle short-drama, gaming, and e-commerce clients' large-scale content generation through to ad delivery, ultimately results-oriented and results-empowering. While producing large volumes of content — whether short-drama content or ad/brand creative materials, including intelligent delivery — we have many insights. I'm here to learn and discuss, thank you.
Jiang Zhiqiang: Thanks, Yin Zong. Next, Meowster — I don't know if I'm pronouncing it right, is it Meowster?
Nie Yonggang: Yes, it's Meowster. Hello everyone, I'm Sawyer, product lead at Meowster. The name comes from wanting to make an AI cat with a sense of life, vertical in the AI emotional companionship track. What makes us different is we don't want humanoid; we want to use a cat form, because of a cat's uncertainty — we want that. Human-to-human pressure is very high; the core is the Japanese market, where emotional pressure is even greater. We want to make cat-like experiences where you can't guess what it'll do next, or it might surprise you, or do something annoying, but it'll be interesting. That's basically it, thank you.
Jiang Zhiqiang: Thanks, Nie Zong. Actually, the introductions were a bit short — let me have you quickly supplement: city base, is the tool or service domestic, global, or both? I know you're in Beijing, Wang Zong is also in Beijing, Yin Zong is in Guangzhou right? Nie Zong, which city?
Nie Yonggang: I'm in Shanghai now; the founder and operations team are in Tokyo, Japan.
Jiang Zhiqiang: I've had a chance to learn about several of your products. Whether you're doing domestic or going global, please quickly supplement.
Liang Wei: Let me quickly supplement. Recently I used this — because I come from film, I recently worked with the Chopstick Brothers on their MV, and "Xiaoyao Xian" is quite popular online now, roughly 200 million views. Actually, why do we say we're committed to bringing AI multimodal capability to FDE service? The core is that today's multimodal generation capability must ultimately land in application scenarios, truly helping traditional film and TV reduce costs and increase efficiency, and helping everyone create and deliver truly consumer-grade content. So this year we spent a long time launching the pro version. Previously global, currently 1.5 million users across 174 countries. MovieFlow is at movieflow.ai — everyone can check it on their phone. Since the Studio version launched in May, it has served over 100 film and TV crews including Hollywood, Chinese, Japanese, and Korean. Everyone uses our Studio for pre-rough-cut and pre-shooting services for traditional film and TV, and to redesign the entire production pipeline for series and film shooting; some enterprise clients also use us for asset management.
So I think today talking about Agent — from my perspective, AI video Agent doesn't exist, video Agent is dead, Skills have no value either. We'll discuss this in a bit.
Jiang Zhiqiang: Wang Zong, please supplement. I just looked at your app — it's interesting, all in English, with 10 million gamers. Expand on this?
Wang Bihao: Our business roughly splits into China and globalization. The China URL is Doudou AI; overseas is Hakko AI, Japanese for "box" — AI experience jumps out of the box. Besides Beijing, we also have an office in Shanghai at Morse Space; welcome everyone to come chat. The main product model: on Doudou, we mainly cooperate with Intel, AMD, NVIDIA, Lenovo, HP and other PC integrators; the software experience is integrated into shipping AI PCs — so if you buy an AI PC, you'll see Doudou pre-installed. OK. Key overseas markets are T1 countries — North America, Japan, and Korea. Why? Because those countries have stronger payment ability. To put it plainly, we want to make money.
Jiang Zhiqiang: Yin Zong, anything to add? Tools already launched, tell us a bit.
Yin Tianming: We're different from many AI tool companies; our team is business-driven, essentially business-driven — we come from operations. I did 20 years of gaming, previously VP of Product at 37 Interactive Entertainment, then two years of short-drama production and distribution. Entering AI became an AI technology product company essentially because we had business needs — deep ad-buying needs and experience driving our own transformation into a tech company. So on the product side, we combine our own business needs, distilling accumulated algorithms, know-how, and different vertical models — like gaming ROI model algorithms, short-drama content generation for improved quality, plot scripts — all distilled into hit-rate algorithms. So we may have some differentiation. On current usage by short-drama, gaming, and e-commerce clients, a few numbers: of short dramas we've co-produced, an average of 5–6 out of 10 are hits — roughly 50% hit rate. Ad creative material pass-through rate is 30%, versus under 5% the traditional way. Average customer acquisition cost reduced 20%–40%, ROI improved 15%–20%. These numbers may have clear appeal for final ROI conversion. We're more about empowering AI products and technology for end results.
Jiang Zhiqiang: Whatever product you make — content or short dramas — the most important环节 is distribution and marketing. Nie Zong, your app is launched. Tell us, is it only Japanese?
Nie Yonggang: We made Chinese, English, and Japanese versions, launched on the App Store around the beginning of this year, mainly Japanese market, plus U.S., Europe, Hong Kong, Macau — Simplified Chinese not yet. Global users approaching 100,000, core users 93% Japanese. Founder background is IP incubation, previously based in Japan. Choosing a cat, and future ideas, is about grafting onto — not just AI companionship but integrating a lot of gamified content. Including a concept now called "Cat Town" — cats in the town are self-made, with their own elections, careers, lives, and NPCs, people's cats from around the world but no user accounts. I think it's the relationship between individual and group, with more fish-keeping, gamified things. Commercialization has started, mainly subscriptions and gamified items; later may expand to offline, currently mainly APP scenarios.
The Real Value of Agents: Gimmick or Industrialization?
Jiang Zhiqiang: As you just heard, from a token-consumption angle — if we discuss commercialization in a bit, token consumption is a very important indicator of whether a product's business works. The industries you're in are either close to high-growth short dramas — globally growing, China leading, AI penetration rising fast — or AI entertainment, or companionship; all high-stickiness, high-token-consumption scenarios.
I want to ask you all a question. Liang Zong just said he thinks AI film/TV Agents aren't useful; others may have different views. Thanks for that provocative opening. Let me state my view first: many Agents aren't To C; behind you there may be many Agents working in tool flows, and the smarter they are, the simpler the product can be. Let me open with that — so starting from this provocative statement, you guests: do you think the real value of Agents is a gimmick, or have you already seen industrialization improving? Also, the team I incubated with a listed film company over the past half year has many production steps; I've seen some steps where AI plays a bigger role, even has automation potential, and reaches premium quality. But I want to ask you, based on your industry observation or frontline experience: have Agents really reached industrialization? Or is it the paradox that Agents have no value?
Nie Yonggang: Our thinking is that the underlying layer is also Agent mode. I don't think Agent is an empty shell; the underlying layer has many settings, like solid-md, lots of underlying content, including OpenClaw which was very popular in February and March. The Agent concept and underlying logic make complete sense. I like using Jarvis as an example of Agent — everyone wants that future state. I personally think it won't be too short or too long; everyone's own Jarvis will arrive soon, just with different ways of operating and using it. From a company perspective, Agents mainly serve internal company or user ends — two sides. If individual, more like a private assistant/butler managing all specific affairs. I've built many Agents myself to manage work items, schedules, process management, report generation, management content — all can work well. Including now with Codex 5.6, the effects and efficiency gains are really fast-paced. So the Agent concept definitely has practical value.
Jiang Zhiqiang: Your Meowster — I haven't had a chance to use it yet — is the content pre-rendered and just animated by AI, or can some content be naturally and automatically generated based on user interaction?
Nie Yonggang: Content is basically all influenced by user behavior, and all content is AI-generated — the cat's own Agent generates it. For example, as interaction stickiness with users grows, the cat may increasingly like them, or may probabilistically still not. If you ignore it for a week or two, there's a 70–80% chance it won't want to interact either, because it feels you can't give it care or sticky companionship. Many things need better feedback based on user input — including feeding the cat, playing with toys, all immediate feedback is given in the moment, based on long dialogue. Starting with dialogue as the core, later it gradually extends to the entire companionship scenario, main space, mall system — the whole integrated content. Each cat feedback is immediate, AI-generated in real time. Based on long-term usage habits, frequency, and communication patterns, the cat knows the user is busy at 7 or 8 PM — writing — and can't reply at 8 or 9 AM, busy with work. This can provide great user feedback. We've been lucky to interview many deep users and chat users, and feedback is that the cat's insights are just right.
Jiang Zhiqiang: Next, Yin Zong — you mentioned earlier your co-founder has rich 37 Interactive Entertainment experience. Earlier discussion highlighted a major selling point: using ad creative material data, with data flowing back to creative generation. Although still in internal development, maybe only internal beta — can you expand? Did this process use Agents and Skills?
Yin Tianming: I was VP of Product at 37 Interactive Entertainment, responsible for overall product architecture, planning, strategy, and对接 with many streaming platforms. The other two co-founders: one spent over a decade at 37 in page-game and mobile-game ad buying and distribution, handling tens of billions in ad spend. I've also done many gaming products from billions to tens of billions. The other co-founder spent over a decade at Tencent as distribution lead, understanding the entire operations process and big-company tech middle-platform system, and has done many hit products from billions to tens of billions. The three of us are an iron triangle — business iron triangle — across product, operations, ad buying, and distribution/marketing.
Because we're business-type, practice-born — essentially, Jiang Zong, I'm interested whether your rich experience has already been distilled — yes, so I was pointing to that Agent — I know, that's why I mentioned the predecessor. Our team is strong with lots of product practice and has built many revenue-generating products, but it's unrealistic for the team to build many products; there are business needs: producing large volumes of creative material, operations data analysis, attribution, and re-delivery. Wanting business-driven technology — can we make it an Agent, a digital human clone? You can think of it as making the three of us into digital human clones. Essentially, an Agent is a digital employee, but if you just call it an Agent, it may be generic, not understanding vertical-domain application — how to do it, how to understand product data, how to attribute and iterate — not professional, maybe just a basic employee. How to teach it? First step: distill past experience, data understanding, product understanding, and information into it. Second step: integrate algorithms so it knows how to work, understanding different products and types — for example, dramas split into male/female frequency, cultivation/transmigration and how to make creative material; games in different categories MMO, card, SLG — feed it, distill experience, plus algorithms, then it calculates correspondingly; otherwise it's chaos. What we call algorithm now may in the future be called a world model — essentially giving the Agent emotion, vertical-domain Agent emotion, to assist judgment. This is the source of our overall product thinking. During distillation — the data changes mentioned — after distillation, integrated algorithms produce actual applications in short dramas, gaming, and e-commerce, giving ToB clients real ROI-oriented feedback.
Jiang Zhiqiang: Liang Zong, explain why you think AI film/TV Agents are uninteresting?
Liang Wei: Mainly because recently lots of WeChat articles started pushing AI video Agents and Skills. Honestly, this is like last year's GPT and Ghibli style — burned for a few days and over. There are many video tool flows and workflows now, and tons of Wong Kar-wai director-style Skills, director-style Skills — to people who do professional content and come from film, it's funny. First, if good content could be produced this simply, it would be easy; there's no necessary connection between a good tool and a good work. What Agents should do is personalization — large models can, generic Agents empowering each creator makes sense in the tool. But piling up a bunch of director styles — did they get the director's copyright? On what basis do they use Wong Kar-wai's name? On what basis? Should we distill Wong Kar-wai? On what basis? What you distill is just a shell — you distill the visual feel of Wong Kar-wai's films, but what does that have to do with storytelling? So at the professional level, Agents make more sense in workbenches and tools — helping users create better within the creation and content needs process. That kind of Agent makes sense. But generic video-related Agents — whoever announces Agent capability or Skill capability — completely meaningless.
Jiang Zhiqiang: But if a tool like this now takes a script, with one or more Agents behind it, automatically decomposing storyboards per the director's detailed requirements, designing assets, and auto-generating shot prompts — is that useful? Regardless of whether it's called an Agent.
Liang Wei: What you described is actually the first step of our pro version now — the script auto-parsing process does exactly that. But after parsing, it doesn't mean everyone can use it to make good content; you need to adjust and refine on that basis.
Jiang Zhiqiang: At least efficiency improvement, right?
Liang Wei: Efficiency improvement is certain; every Agent helps everyone improve efficiency.
Jiang Zhiqiang: Wang Zong, do you have a different view? Share from your product and experience.
Wang Bihao: I'm also interested. Around this topic — first, what is an Agent? Why is it called an Agent? An Agent has autonomous decision-making capability. Why? The key is the loop: after receiving a task, decompose it to various professional skills, synthesize, the Orchestrator combines them, and finally a reviewer reviews the work, with a feedback system telling you how well it was done. In my view, right now there's only one type of Agent that works: the coding Agent. All other Agents don't actually hold up. Why? Because there's no reviewer — no one can judge whether this Agent did a good job or not. Coding works because there's a standard: whether code is correct and runs can be clearly told. Whether a film is good, whether a video is well-made — only the audience can tell you, and it's subjective. So temporarily no other domain's Agent works.
Another topic: a friend recently told me he feels all Agents can't be done — there are three things Agents can't do, meaning demand doesn't exist. Why? Users fall into three categories: first, no ideas — give them an Agent and they don't know what to do; second, have ideas but can't write prompts well — prompt decomposition and writing is complex; third, have both ideas and prompt-writing ability — they don't use this product, they go to Claude Code and Codex, why use yours? So all Agents are losing their user base. That's his view, that Agents can't be done. But right within it lies the next opportunity: people with no ideas — because the barrier to ideas is high, requiring natural-language input — if we could feedback ideas through light interactions like clicks, swipes, likes, comments, would that work? Second, people who can't write prompts well — this is the biggest escalation problem for Agent entrepreneurship: how to bring Claude Code to ordinary consumers, because Claude Code is a magical product experience. I don't have the answer now, but this is a very certain and reliable direction — if someone can use Claude Code to make video, it definitely works. Third, those who go straight to Claude Code without a product — that also contains opportunity, because building loops in different domains isn't all using MCP and Agent Core systems; there are business-integration secrets that require true domain experts to participate to uncover that domain Agent's truth. That's some of my views.
Jiang Zhiqiang: I want to follow up — how many Agents do you use in your product now?
Wang Bihao: I'd say quite a lot. I'm a heavy Claude Code and Codex user, averaging about 300–500 commits to projects per day — those are our main Agents. In the product, we build Agents for business flow: data analysis, user feedback, all implemented through Agents. The most important thing about an Agent is evolutionary — it changes; with use, the Agent rewrites its own code. That's the most important point. In actual user delivery, we also continuously implement reviewers — when users generate video, game, or text that needs review, the Agent reviews it while the user also reviews it; likes and feedback are collected in real time into a judgment system that decides OK or not, and if not OK it goes back for rework, next iteration. The biggest advantage of this Agent loop over traditional software development is real-time evolution. Before, software received 10 user feedback items, iterated a new version next week, then collected more feedback. Now it's not needed — users give real-time feedback and the software changes in real time. That's what I consider a true Agent system.
Jiang Zhiqiang: Agent loop is a hot topic in the AI industry now. Next, let me warmly introduce our fifth guest.
Luo Yuan (Shanzhu): I'm Luo Yuan, Executive Director of Huanjing Jiyuan, also a self-media creator and AI video super-creator, a super-creator at the neighboring Pixworks. Briefly introducing myself.
Jiang Zhiqiang: I'll follow up with you, related to the topic. As a super-creator, you must have touched many tools; before new tools come out, big-company tools and startup tools all ask you to test. I don't know what type of content you mainly create now? From the highest-frequency creation content, among the endlessly emerging tools — including big companies making tools, the two CEOs here are making tools — where do you find tools useful? What pain points remain unsolved? You're most qualified to speak to this.
Luo Yuan (Shanzhu): Currently I find Tapnow and Pixverse most usable; their generation platform operations, chained operations, and coordination are relatively mature, with attention to small operational details, speeding up creation. For example, you can click an image on the platform and directly @ it for very fine-grained operations, and when connecting, you can copy a whole string. There's team collaboration, very convenient; I hope major platforms also consider adding this — it's very important to us.
Jiang Zhiqiang: We don't intend to advertise these tools. But I want to challenge this, given my half-year incubating premium short-drama self-production plus Agent team. Observing AI gacha operators or AI production staff: the pain point is that if it's not self-produced drama but serving platforms, platforms want volume AND premium quality, which may not be enough. Content must be both premium and efficient; the canvas experience may not suit everyone. From a "want both, want three" perspective, what can tool makers — big company or startup — still improve?
Luo Yuan (Shanzhu): You could add, for example, converting scripts into prompts suitable for Jimeng or Kling video generation models, doing initial conversion. Qinghe Tech's platform is also okay; they're doing it but maybe not quite perfected, on some fine operations. Because we're in the premium track, maybe not huge volume, trying to keep it few but refined.
Jiang Zhiqiang: I'll follow up. For premium content, whether on projects or personal work, what Agent do you start using? Do you use Claude Code for efficiency?
Luo Yuan (Shanzhu): From a creator's video creation perspective, I don't actually use video Agents much, because of a series of problems. Communicating with AI is laborious; most of the time using an Agent is like communicating with an employee who may not understand you, increasing communication costs. Plus, using an Agent — now GPT's conversational content generation causes secondary iteration images to break up and get noisy, requiring rework. So I don't let my employees use Agents; it increases communication costs.
Jiang Zhiqiang: Next, I want to ask the two CEOs making tools. Big companies are also entering tools now; although big-company products may not all be deep enough, leaving room for startups. But I especially want to ask Liang Zong and Yin Zong: as big companies enter tools, what are the challenges and opportunities?
Liang Wei: First, let me承接. MovieFlow going forward is no longer just a tool but a global AI film/imagery platform — from workstation to distribution. After the APP launches, it'll also be a horizontal-screen premium AI global content platform, playback platform, viewing platform. We're also investing in many frontline film directors and companies to co-create horizontal-screen premium AI imagery content. For me, multimodal AI video capability is like reconstructing and redoing the entire film and entertainment industry. Going forward, everyone can use many celebrity digital assets in the MovieFlow APP — this is unique, after all, I'm originally from film and entertainment. Recently Chopstick Brothers' Xiao Yang and Wang Taili's likenesses have been entering the asset library; other actors are gradually coming. In the future, ordinary creators can commercially use film-level asset materials for creation, truly doing this.
Big companies each have their tracks; for the entire AI imagery space, like Tencent, Youku, iQIYI, will the future produce the most AI track? Seeing AI video suddenly in iQIYI — is it real footage or AI content? It's strange, especially as deepfake multimodal generation quality keeps improving, indistinguishable from real filming or AI-assisted/multimodal direct generation. AI imagery needs an independent broadcast platform separate from traditional film/TV production; having that is a completely pure blue ocean. It's certainly not just the AI short dramas and comic dramas that ByteDance pushes every day. AI video capability in comic dramas and short dramas has only exposed 0.5%–1%; there's still 99% that can produce large numbers of BBC documentary films and better films with better production. The tool layer is already very simple; anyone can hand-build a workflow in a day or two, but whether the workflow has accumulated creative assets and value matters. The best tool doesn't mean giving you an IMAX camera and you shoot like Nolan; there's no necessary connection between a good tool and a good work.
We're working to connect traditional film production resources with the current generation — students entering film school today, directing as freshmen, in four years there won't be that many films to shoot, but today they can normally use AI for film-level creation. With requirements for story, performance, and creation, with a baseline — that's AI video. Last year at the Unique conference I said AI video should be called Aideo, as a new field and industry, like film vs. cinema and digital shooting producing television and phones producing short video. AI multimodal should produce Aideo, with a new audiovisual language system and consumption system. As MovieFlow, we hope the industry sees this, pushing professional filmmakers and AI creators to co-build the ecosystem, truly making money and commercial consumption — then it holds.
Yin Tianming: Liang Zong said it well, basically many consensus points. Let me supplement: in the future, more based on big companies, big companies more serving pan-users, not very vertical; many vertical pan-users are mostly To C, vertical more facing ToB clients or users. Most ToB clients are small and medium, mid-tail and long-tail, not giant ToB film companies; they need empowerment. First, they need to know how to make good content, use good tools or AI video content generation platforms; algorithms and accumulated digital assets can empower them to produce content with hit potential. Also, I can help you sell it, help you make money — no need to对接 platforms or streaming yourself, no need to do ad buying yourself; the tool can help you. Ultimately it must convert to revenue. Future big-company differentiation: big companies won't help you do these things; you need these to continuously empower.
Jiang Zhiqiang: Finally, Wang Zong, please supplement. Earlier we discussed AI film content getting more refined, and the problem to break through is so-called consistency, especially scene consistency; current video generation model architecture lacks a z-axis. Although what you do is interesting with large gamer communities, from the next breakthrough point — can you briefly expand on what you're doing in this area?
Wang Bihao: I believe Shanzhu一定 has a pain point: why can we only generate 15-second clips? 15-second clips, wanting 4K means 4–5 seconds. Why does this happen? It's interesting: this generation of models — CogVideo, Kling, Wan — all based on DiT architecture, Diffusion Transformer. Diffusion is the diffusion model, each frame gradually diffused from pixels, deciding what the picture looks like; Transformer is time series. Through large amounts of video training, character consistency across frames is maintained during diffusion — that's the basic principle. Diffusion volume and time-series volume have extreme requirements for parameters and compute. Current mainstream model parameters are 20B to 30B; either you break through parameter count with more compute, but architecture has hit a bottleneck.
Why bottleneck? The root is researching VLM — how to use visual recognition to understand what video is saying. Researching VLM mainly to achieve game screen recognition, then found that training world models is essentially this. Training a 20B to 30B parameter world model requires 1 billion videos; data is extremely scarce. Liang Zong, in the film industry, knows a director who dares to shoot a 10-minute single take is very bold. Training consistency models requires 10+ minute long shots, letting AI understand no cut, no perspective switch, so it can understand why current AI video generation — someone opens a door, walks in, walks out, and the person, clothes, and environment change — because in the training corpus, when a film/TV door opens and someone enters, the perspective and lighting all change, so AI thinks the world is just like that, and that's how it understands the world.
From this angle, the important work we're now doing is collecting large long video segments from visual data — real-world directors don't shoot long takes, but game worlds are all long takes, first-person perspective directly producing an hour of continuous footage, with lots of corpus, but still insufficient for training — roughly needing 100 million videos. On one hand, trying to collect more data; on the other, doing small-parameter experiments, provided in the new product Color. The core concept is users' real-time input influences short-drama generation direction — the short drama starts with a domineering CEO abandoning me, and the next generation is written by the user in real time. The most important generation requirement is ensuring the domineering CEO doesn't suddenly change from square-faced to round-faced. This is some of the results from our recent long-shot data collection; briefly introducing these.
Jiang Zhiqiang: The conference is signaling we're over time. That ends today's forum. One last thing: I'm currently Venture Partner at Heli Capital; if you know teams in the AI circle that are doing well with products showing some data, feel free to recommend them to me.