跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

From a 200-Person Crew to 20: After AI Video Kills Gacha Pulls, the Most Expensive Role Surfaces

What you hear today is not what you actually ship.

Last year, 20% of MovieFlow’s code was still written by humans. This year that number is zero — 100% generated by AI.

The same team is compressing a traditional 200-to-300-person film crew down to 20.

Put those two numbers together and you get the real water level of AI video right now: machines are eating execution-layer jobs faster than anyone expected.

Over the past year, “gacha” — rolling the dice until one random generation looks good — has dominated the conversation. Getting one great clip by luck is easy. Producing one reliably, every time, is the hard part.

In a recent Unique Capital roundtable, four companies attacking that problem from different directions sat down together: Jifan Liuxing (3D and physics-controllable video generation), MovieFlow (video agent platform), Anigo (comic drama and AI audiovisual), and Xuanjia Technology (industrialized mass production). After the full conversation, my sense is this: the “gacha” headache is being engineered away fast. But what it exposes underneath is more interesting than gacha itself.

Character consistency is the worst gacha offender. Making a character look good in a single clip is not hard. Hard is making the person in shot 15 the same as in shots 2 and 3.

The four companies’ solutions barely overlap.

Xuanjia CEO Qin Lin’s answer is the most “engineering.” His judgment: as long as foundation models use the Transformer architecture, denoising will always produce per-generation character variance, and the model side alone cannot fully solve it.

So they wrap three engineering constraints outside the model: sample anchoring when the lead character enters the pipeline, automatic redraw when variance is too large; dual libraries of character and scene assets; and prompt engineering using professional cinematography and directing vocabulary.

The result is quantifiable: 95% mechanized production, 5% human tuning. Their KinoAuto module takes a novel as input and, in four hours, outputs a film with near-zero human intervention. A 120-minute short drama costs 5,000–10,000 RMB and one person finishes it in a day.

MovieFlow founder Liang Wei’s path is the opposite: real actors as the base. They came from film. Today they build a library of real performers. A real person exists in the physical world; once captured clearly and weighted into training, calling them is like calling a photo or asset, with essentially no consistency problem.

The market response is strong: in nearly a year since launch, MovieFlow has reached 1.5 million global users across 170+ countries.

Jifan Liuxing’s Huang Naiyuan is the most candid. Asked how they solve character consistency, he says directly: “We still leave it to the foundation model.” They are all-in on 3D and physics-controllable generation — keeping backgrounds, objects, spatial relationships, and camera motion physically coherent. For example, fix the scene structure and camera position first, then hand off to the video model, improving cross-shot consistency in the same scene.

Anigo’s Wang Jun bets on agent engineering: which task goes to which agent determines how stable the character stays.

Four paths, different principles, each delivering. That itself says something: character consistency has been downgraded from a research problem to an engineering selection problem. Pick any path — they all get there.

Dig deeper and “the face stays consistent” is the easiest layer.

Wang Jun splits consistency into two layers: visual and narrative. At the visual layer, foundation models are already good. The hard part is long-form: single generations have a roughly 30-second boundary. Beyond that, the problem becomes how agents coordinate to maintain the character’s narrative coherence.

His example is vivid: in one shot the character is euphoric, in the next suddenly reserved. As Wang Jun puts it: “At the start he might be a lunatic.”

Liang Wei calls this performance consistency and says it more bluntly. We watch acting; a face that merely looks alike is meaningless. In a single scene, a character may move from anger to calm, back to anger, and finally to reconciliation. That flow of states is the scene. Pure AI doing that today is genuinely hard.

Qin Lin also admits the boundary. The remaining 5% clusters in extreme situations, like heavy backlight. But he offers a useful counterpoint: in real shooting, backlight also makes faces unreadable. Some of AI’s weaknesses are just photography’s weaknesses.

The conclusion around camera control has little to do with technology.

Wang Jun says people often tell AI “push the camera in a bit.” What does “a bit” mean? A real first assistant director would interrogate the request until it is extremely specific: how to move, how far. It is an extraordinarily vague instruction.

Liang Wei names the issue outright. Anyone today can sit in a director’s chair and yell “Action.” AI is the same: you state the goal, yell Action, and a crew of hundreds starts working — actors perform, cameras roll.

The hard part is yelling “Cut.”

“You yell Cut, the whole set stops and looks at you. They are waiting for another result: was this take OK or not?”

Wang Jun adds a line anyone who has used AI generation will recognize: typical users only say, “I feel something is off, think again.”

What counts as good? Liang Wei’s judgment: large language models handle the middle 98% of the work, but everyone is waiting for that final directorial judgment.

So the tool layer has no essential difference: “Today you have a 3D director console, tomorrow I will too.” The real difference is the user. Give Zhang Yimou the same scene and it comes out differently.

AI raises everyone’s floor from 0 to 80 or 90. The result: directing skill becomes more valuable, not less.

When a crew shrinks from 200 to 20, who stays?

Liang Wei’s list: screenwriter, director, an art director fused with DP capability, a performance director who understands acting and story, the principal actors in the physical world, and editor.

Editing gets called out specifically. Short dramas and comic dramas can be cut in one click because 2s, 3s, 8s, 10s shots have relative standards. Film editing has no standard.

Montage can cut abruptly for a shock; or deliberately hold the cut, letting you feel impatient and wait for what comes next. Editing is third-order creation. AI cannot replace it today.

Qin Lin gives the same conclusion from another angle. He says two things agents and models cannot solve.

First, the script. Many people use AI to write novels and scripts, but none touch the soul of human emotion, because “machines can only go from the known to the known, not from the known to the unknown” — and humans’ greatest strength is imagining unknown worlds.

Second, a top director’s grasp of framing and aesthetics. Taste is hard to parameterize or distill. Without human intervention, pure AI’s ceiling is around 70–80 points.

The future collaboration frame is clear. Liang Wei’s version: humans own the first 1%, the 1% after delivery, and the 1% of next-round direction; the middle 98% of each execution goes to AI.

What the 20-person crew loses is large amounts of execution physical labor. The people who create and exercise judgment? Not one leaves.

There is one unavoidable judgment: is the next Vibe Coding moment in AI products video generation?

Qin Lin breaks down why Vibe Coding worked: code has conventions, standards, high repetition frequency. The more standard and high-frequency, the easier to commercialize.

By that logic, the only AI-video scenario truly working at scale today is the short drama — because short dramas match code’s logic: corporate, process-driven, standardized, low creative share, classic popcorn content.

His future video world splits in two. One side: premium film. A great director plus good tools can make what used to cost hundreds of millions for a few million to ten million RMB. The other side: purely industrial production of short dramas, ads, e-commerce, digital humans. Combined, 10–15 years out, a 10-trillion-RMB industry.

Huang Naiyuan’s caution is cooler: Vibe Coding will likely solve only a small part of video generation. When building products, see clearly what foundation models will not cover in the near term. When the tide rises, simple processing products drown first.

Liang Wei’s answer is sharpest, and the best advice for everyone. First, don’t be anxious about technology. Consistency and crew-management tools will be built; creators don’t need to master each one. Second, enter now. Technology has progressed beyond imagination. Mature series they are producing with Hong Kong filmmakers already show no visible AI traces.

He even predicts an “African Story of Yanxi Palace” with an all-Chinese production team, because the world’s strongest AI video production sits in China.

His last line is worth closing with:

What you hear today is not what you actually ship.

Originally published by Unique Research on Unique Research Substack on September 24, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research