---
title: "74 Episodes in a Week: He Says the Most Expensive Cost of AI Video Isn't Compute at All"
author: "Unique Research"
sourcePublication: "Unique Research Substack"
originalPublishedAt: "2026-09-20T14:19:29+00:00"
canonical: "https://ffcap.cn/en/research/74-episodes-in-a-week-he-says-the"
source: "https://uniqueresearch.substack.com/p/74-episodes-in-a-week-he-says-the"
language: "en"
---

# 74 Episodes in a Week: He Says the Most Expensive Cost of AI Video Isn't Compute at All

Ultimately, the model determines what we can do, the creator determines what we make, and the market determines what stays.

A 74-episode Cambodian short drama — from character modeling, shot generation, compositing to Khmer voiceover — was completed in seven days.

The actors only existed as a few photos.

This is _Where You Belong_, a recent overseas project by Xuanjia Technology. In the past, organizing a separate shoot for a small-language market like Cambodia would have been hard to justify commercially.

Now, a small team plus an AI production pipeline made it happen.

Xuanjia was founded in 2013, starting in smart audiovisual technology, with an existing audiovisual business covering millions of household users. Logically, continuing as a stable content technology company would have been a comfortable path.

But in recent years, they’ve bet on the heaviest layer: their own video foundation model Kino, while also building a creation platform, content production, industrial bases, and global distribution. They’ve now produced over **100,000 minutes** of AI content.

But models get replaced every six months. If a stronger general model appears tomorrow, what’s left of your own model’s value?

Qin Lin, founder of Xuanjia, doesn’t assume any model will stay ahead forever. He’s matter-of-fact: video generation is still evolving fast, and a stronger general model emerging at some point is perfectly normal.

What Xuanjia really wants to accumulate isn’t a set of static parameters, but four categories of continuously growing assets:

-   High-quality industry data
    
-   Workflows validated by real projects
    
-   A team that understands content production
    
-   A feedback loop from production to distribution
    

External models getting stronger isn’t necessarily a threat to Xuanjia — it might actually expand the entire industry’s capability boundary.

They can both absorb and combine external capabilities, and continue training and optimizing Kino for industry scenarios.

> “The ultimate competition isn’t about one model benchmark ranking. It’s about who can more stably convert technology into qualified content, and control quality, schedule, and cost.”

That’s where the conversation’s underlying logic diverges.

If you think AI video competition is about model scores, then building your own model looks like a gamble. But if you think competition is about the ability to consistently deliver content, then the model is just the entry point — the real moat is the entire production line behind it.

Of course, a moat needs a grip. General models are so strong — why does Xuanjia still train Kino?

Qin Lin says general models are important, and they’ll stay open to using the best capabilities in the industry where appropriate. But there’s still a long distance between “can generate beautiful video” and “can stably produce film and television content.”

Industry projects care about more than just how stunning a single shot looks:

-   Can characters stay consistent across shots?
    
-   Can costumes and scenes remain continuous?
    
-   Does the shot serve the narrative purpose?
    
-   Is batch generation controllable?
    
-   Can edits be precise to a local area?
    
-   Do cost, speed, copyright, and data security meet delivery requirements?
    

> “Kino’s value isn’t to replicate a general model. It’s to connect model capabilities with script analysis, character assets, storyboard design, shot scheduling, sound production, and post-production compositing — specifically for film, short drama, and comic drama production.”

I pressed him: what does “understanding film” or “understanding short drama” actually mean? Is it cinematography, character consistency, or narrative understanding?

“All of these matter, but real ‘understanding film’ isn’t any single metric — it’s understanding the relationships between these elements.”

A shot’s quality can’t be judged out of context. The model needs to know what this shot is saying, what emotion the character is in, how adjacent shots connect, and whether this shot is establishing information, creating conflict, or delivering emotional release.

Short drama adds another layer: how to quickly establish conflict in the opening, how each episode forms a hook, how characters and scenes can be reused at low cost.

> “What creators need isn’t one-off lottery-style generation. It’s characters, scenes, actions, expressions, camera movement, and sound that can all be controlled, modified, and reused.”

“Lottery-style generation” is a vivid phrase. Win and you post it to social media; lose and you’ve burned compute.

Industrial production wants controllable, continuous, modifiable output. Whether the image is beautiful comes second.

Talking about industrialization, Qin Lin gave a clean definition:

> “Industrialization isn’t generating a lot of video. It’s being able to consistently deliver under defined time, budget, and quality standards.”

There’s a chain of steps in between:

-   Scripts need to convert into executable shot plans
    
-   Characters and scenes need asset management
    
-   Generated results need to stay continuous
    
-   Failed footage needs quick diagnosis and redo
    
-   Different models and human steps need to coordinate
    
-   Finally, dubbing, music, editing, review, and multi-platform delivery
    

> “A demo can show only the best ten seconds. Industrial production has to handle all the substandard results.”

I suggest everyone who’s been wowed by AI video demos read that three times.

What you see at a launch event is always those ten seconds. The rest — mismatched lip sync, faces that changed — that’s the daily reality of production.

There’s an industry number rarely talked about openly: **effective footage rate**. Generate 100 minutes of footage — how much actually makes it into the final cut?

Qin Lin didn’t give me a pretty number. He says the ratio varies widely by content type, visual complexity, and quality standard. Continuous narrative content has much higher requirements, because each shot must not only stand alone but also stay consistent with adjacent shots.

Currently, rework concentrates on these areas:

-   Multi-person interactions
    
-   Complex physical movements
    
-   Character lip sync and dialogue performance
    
-   Continuous scheduling
    
-   Character consistency across different shot sizes and angles
    

But what he really cares about is the **predictability** of rework. Can you know why it failed, and through character assets, prompt templates, shot breakdown, model scheduling, and human correction, steadily raise the effective rate?

> “Rework isn’t scary. Unpredictable rework is the biggest problem.”

Moving from production to data, I asked a heavily debated question: is the truly scarce data of the future more online video, or high-quality correspondence between scripts, storyboards, shots, and final cuts?

Qin Lin’s answer is clear. Public web video helps models learn world knowledge, visual rules, and basic motion. “But what industry-grade video models most lack isn’t raw video quantity — it’s structured, interpretable data that maps to production workflows.”

He gave an example. A work goes from script to character design, storyboard, shots, performance, sound, editing, and final cut — why each decision was made, what changed between revisions, and what result the director and market ultimately accepted.

That kind of correspondence is far more valuable than an isolated final cut.

Two types of data will be most scarce:

-   High-quality audiovisual data with clear copyright and stable quality
    
-   Process data and feedback data from real production
    

> “The first tells the model what good content is. The second tells the model how good content is made and revised.”

In other words, final cuts teach the model what’s good. Process data teaches the model how to get better.

The second type isn’t on any public network. It’s only in companies that actually produce.

I posed a question to Qin Lin: if a team today takes **5 million yuan** to enter the AI short drama or comic drama industry, what are the three easiest miscalculations?

His answer is almost a pitfall-avoidance guide.

Many people only calculate generation fees, without accounting for filtering, rework, editing, sound, review, and project management. They also underestimate the time needed to build a stable workflow.

> “Low unit price doesn’t mean low total cost.”

Content being made doesn’t mean it gets seen. Traffic costs, platform revenue share, payment cycles, and operational capability across different markets often affect final profit more than production cost.

Business plans easily assume every project achieves average returns, but content industry revenue is typically highly concentrated.

What really needs calculating is: how many of ten projects break even? Can failed projects’ costs be controlled? Do successful projects cover the whole portfolio?

> “The most dangerous thing is directly treating model demo performance as stable capacity, then treating stable capacity as market revenue.”

Two layers of “treating as,” two traps. A good demo doesn’t mean stable capacity; stable capacity doesn’t mean people will pay.

Following up, I asked how he sees the gap between “can produce” and “can make a hit.”

He says these are two completely different capabilities. Producing solves delivery. Making a hit involves user insight, emotional resonance, rhythm design, distribution timing, and some randomness.

AI can analyze historical content’s themes, character relationships, pacing, user comments, and conversion data — helping teams screen directions, simulate plots, and quickly make multiple versions for testing.

> “But AI provides probabilities and leads, not hit guarantees.”

Generation costs keep falling — that’s industry consensus. But Qin Lin points out a counterintuitive direction: costs don’t disappear, they transfer.

Model compute is still an important cost. But the more widespread generation capability becomes, the cheaper it is to simply get an image. Value moves toward scarcer links. What will be truly expensive in the future might be:

-   Good IP, good scripts
    
-   Directing and aesthetic judgment
    
-   Copyright compliance
    
-   User acquisition
    
-   Global distribution capability
    

> “When everyone can produce, what the market lacks isn’t content quantity — it’s content worth users’ time, and the ability to deliver that content accurately to the right audience.”

On the 2026 AI short drama boom, his judgment stays half-cautious.

Real demand exists: user consumption of video content keeps growing. Many vertical and local markets were previously unserved because production costs were too high. AI makes niche themes, multilingual content, and rapid iteration economically viable for the first time.

But supply growth doesn’t equal demand growth. As lots of similar content floods in, user time is still limited, and platforms will quickly raise quality bars.

> “Capacity dividends will gradually fade. Content value won’t.”

To judge whether a company has long-term capability, don’t just look at how many minutes it produced. Look at completion rate, retention, payment, repurchase, and whether the IP can be continuously operated.

Polarization will likely happen: one end is low-cost content monetized through scale, efficiency, and precise distribution; the other end is premium content with creativity, aesthetics, and IP value. In between, lots of vertical content that couldn’t exist before will grow.

But long-term profit doesn’t necessarily belong to the most expensive content. It belongs to content that can form differentiation, reusable capability, and user relationships. Pure low-cost advantages are easily copied.

Back to that **74-episode** Cambodian short drama.

Completed in a week, based on a few actor photos for character modeling, generation, compositing, and Khmer voiceover. Qin Lin says it validated not just generation speed, but the feasibility of a cross-border content production pipeline.

In the past, organizing a separate shoot for a relatively niche language market might not have been commercially viable. AI let them use a smaller team to complete character building, shot production, and local language versions.

But the harder part wasn’t making the images — it was localization.

> “Production capability makes a project possible. Localization capability determines whether it can actually enter the market.”
> 
> “Accurate language doesn’t mean accurate culture. Character relationships, performance habits, forms of address, pacing — even costumes and spaces — can all affect whether local audiences believe the story.”

What’s mature today is translation, dubbing, subtitles, lip sync adaptation, and regenerating some characters, scenes, and visual elements. That already significantly reduces multilingual version production costs.

But true localized creation goes far beyond swapping Chinese for another language. It requires re-understanding local social relationships, values, humor, taboos, aesthetics, and platform habits.

The same conflict can have completely different acceptance across cultures. Even story structure and character motivation may need adjustment.

> “The future isn’t one Chinese content mechanically replicated into dozens of language versions. It’s building a core story asset, then having AI and local creators jointly generate adapted versions for different markets.”
> 
> “AI handles reducing reconstruction cost. Local teams handle cultural judgment. Technical distance is shrinking fast, but cultural distance still needs people to cross.”

That’s also why Xuanjia is building the Kino Ripple distribution platform. Tools can tell you whether a shot generated successfully, but only the market tells you whether audiences want to watch, where they drop off, and why they pay.

Distribution is pulled forward into topic selection and production. Playback, retention, and conversion data feed back into content decisions.

Qin Lin’s judgment: film companies will shift from large crews organized by trade to small core project-centered teams, then connect to model platforms, professional services, and external creators.

Manual demand for basic storyboarding, some asset production, version adaptation, and junior post-production may decline, or transform into composite roles managing multiple AI tools and production steps.

> “But truly excellent screenwriters, directors, art directors, producers, and distribution talent will be more valuable. Because when execution capability becomes universal, what determines a work’s differentiation is judgment.”
> 
> “The people who’ll be more expensive in the future aren’t those who personally perform every operation. They’re those who can build a worldview, define quality, direct people and models, and take responsibility for the final result.”

New roles will also emerge:

-   AI director
    
-   Model and workflow designer
    
-   Character asset manager
    
-   Generation quality lead
    
-   Cross-cultural content operations specialist
    

Roles aren’t simply disappearing. They’re recombining around the new production model.

At the end, I asked Qin Lin: three years from now, what do you want Xuanjia to be understood as? A video foundation model company? An AI content company? Something else?

He said: an AI audiovisual infrastructure company. The video model is the underlying capability. Content production is the validation and application scenario. Global distribution lets the capability truly form market value.

After the interview, his last line kept running through my head:

> “Ultimately, the model determines what we can do, the creator determines what we make, and the market determines what stays.”

The first half of the AI video race is about whose ten seconds is more stunning. The second half is about who can, within defined time and budget, consistently deliver content worth users’ time.

The 100,000 minutes beyond those ten seconds are the real battlefield.

* * *

_Guest: Qin Lin, Founder and CEO of Xuanjia Technology_

[![cover](https://substackcdn.com/image/fetch/$s_!wRYl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F875f8a8b-c2c7-4dad-88e5-44ce70a1cb84_2048x1152.jpeg)](https://substackcdn.com/image/fetch/$s_!wRYl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F875f8a8b-c2c7-4dad-88e5-44ce70a1cb84_2048x1152.jpeg)

---

Original publication: https://uniqueresearch.substack.com/p/74-episodes-in-a-week-he-says-the
On-site reading page: https://ffcap.cn/en/research/74-episodes-in-a-week-he-says-the
