---
title: "After AI Pushes Content to 10 Yuan, the Expensive Part Just Begins"
author: "Unique Research"
sourcePublication: "Unique Research Substack"
originalPublishedAt: "2026-10-01T02:24:41+00:00"
canonical: "https://ffcap.cn/en/research/after-ai-pushes-content-to-10-yuan"
source: "https://uniqueresearch.substack.com/p/after-ai-pushes-content-to-10-yuan"
language: "en"
---

# After AI Pushes Content to 10 Yuan, the Expensive Part Just Begins

[![cover](https://substackcdn.com/image/fetch/$s_!DqkC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98074fd-48ff-4e09-b0d0-c4ba5c1fe194_2730x1536.jpeg)](https://substackcdn.com/image/fetch/$s_!DqkC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98074fd-48ff-4e09-b0d0-c4ba5c1fe194_2730x1536.jpeg)

When one melody can split into 1,800 versions, what is truly scarce is: which version do you want people to remember.

A melody has barely been generated before there may already be 1,800 versions of it circulating for promotion.

A food-delivery worker with no professional training listens to people tell stories, turns those stories into songs, and earns more than 100,000 RMB a month taking orders.

An AI voiceover company less than two years old has 95% of its users overseas, close to 2 million people.

Put these three facts together and the message is clear: the AI-content business is no longer the simple story of “work that used to cost 100 yuan now costs 10.”

At this roundtable during the Extraordinary Awards Chengdu AI Entertainment Conference, Li Zhao (music director of Sichuan Observation), Chen Shenhan (founder of Infinite Picture Book), Song Kaifa (co-founder of VMEG), and Xia Yong (partner at Xingwang Investment) were really answering one question: as generated content gets cheaper and cheaper, what is still expensive?

Li Zhao’s judgment is direct:

“On the surface it’s about raising efficiency; underneath it’s about changing the form.”

The music industry shows this best. In the past, making a song meant going through lyric writing, composition, arrangement, recording, and post-production, with countless rounds of communication and revision in between.

Now many of these steps can be done with one click, even the human voice can be generated. It saves time, of course it saves time.

But the real change isn’t how much labor is saved—it’s that the very thing called “a song” has come loose.

In the past, music was released in fixed forms: singles, albums.

Now a good melody, the moment it enters internet distribution, may already exist in countless versions: different styles, different scenes, different audiences. The distribution path has changed, and so has the generation path.

Chen Shenhan sees the same shift in education. Infinite Picture Book strings together text-to-text, text-to-image, image-to-image, and text-to-video capabilities, then lets children participate in the content production process.

A child is no longer just receiving a story; through plot choices, visual expression, and interaction, the child turns the story into their own experience.

Put simply, the basic unit of content has changed.

Music is no longer just a three-minute finished product; a picture book is no longer a book that ends when you turn to the last page; voiceover is no longer just an audio track pasted on.

AI is pushing content from “something already finished” toward “something that can enter a specific person, a specific scene, a specific emotion.”

That’s why talking only about cost and efficiency feels shallow. 100 yuan becoming 10 yuan is tempting, but it can’t answer the next question: when everyone can produce at 10 yuan, why should users choose you?

Li Zhao tells the story of the delivery worker.

He loves music and has no professional training. He picked a very niche direction: chat with people, listen to their stories, then turn the stories into songs and post them online.

Later, more and more people asked him to write songs—government agencies, scenic spots; people even asked him to write a song to find a lost daughter, or to propose to a girlfriend. He earns over 100,000 RMB a month from orders.

The point of this example isn’t that “a delivery worker can also make music.” It’s that what he’s actually good at was never in music theory.

He understands people; he can find the expressive point in someone else’s story. AI just connects that breath for him.

Tools let different people play to their strengths. But Li Zhao also warns that professional musicians can’t throw away music theory and musical fundamentals, or they’ll lose their taste.

“If everyone generates with one click, music will become more and more homogeneous.”

He gives a concrete example: if you’re making a song about Sanxingdui and AI arranges it with a guzheng, you have to judge whether the guzheng matches the era of Sanxingdui.

That judgment can’t be generated from keywords; it comes from knowledge, taste, and continuous learning.

Chen Shenhan puts it more abstractly, and more sharply:

“Large models provide breadth, width, and a huge pool of content. The key is how a person distills something of their own from it.”

Applied to children: children have rich imaginations but limited expressive ability.

AI can help children dig out what they want to express, making the work part of their own lived experience. The taste, feeling, and logical thinking that grow in this process are what can truly enter their development.

So the division of labor between human and AI increasingly looks like one sentence: AI is responsible for making the pool bigger; the human is responsible for making the choice accurately.

The most expensive skill of the future may be called “picking the right one.”

Song Kaifa does AI voiceover; by rights he should be the one emphasizing technical moats. But the most counterintuitive passage in the roundtable comes from him.

“In the long run, voiceover may not have a technical moat either.”

He says, “There’s a moat in the short term; in the long run it may shrink.”

This carries different weight coming from the co-founder of an AI voiceover company.

His judgment is: as long as there’s enough corpus in a given language, large models will probably be able to do it all eventually.

Right now some low-resource languages, or domestic Tibetan and Uyghur, still lack mature models; but if models cover them in the future, middleman companies may thin out. Users can just have the model generate a video speaking a given language with a given voice.

So what can still be done now?

Song Kaifa names three layers of threshold. First, humans still lead the content. Professional voice actors are expensive and time-limited; you can let other real voice actors perform first, then use AI to correct the shortcomings.

If 80–90% of the emotional expression still comes from a real human, it won’t sound like the AI voice that enunciates every word too clearly.

Second, localization is not translationese. Korean and Vietnamese both have regional differences; clients will ask, which regional accent are you dubbing? If you can’t even tell the language habits of the target market, it’s hard to pull off.

Third, real production is not one person speaking into a microphone. Multi-person dialogue, noisy scenes, seven or eight people talking at once—each needs a unique voice and performance. Single-person voice conversion is one thing; dubbing complex multi-person scenes is another.

It’s precisely because of these thresholds that the market AI truly opens up is more interesting.

Song Kaifa says AI hits low- and mid-tier localization hard, making content that used to be stuck start flowing.

A client in Shanghai told him that some Russian animations suit Chinese audiences but used to be unable to enter the market; now they can be dubbed into Chinese and distributed domestically.

“One Belt, One Road” content heading to Africa—a Swahili version, or a Kazakh version for another region—would be very hard to deliver quickly with traditional human methods.

In one sentence: AI accelerates the flow of existing content.

In the past, a lot of content wasn’t bad; it was stuck on language, cost, distribution, and local accent.

Now these bottlenecks are being cleared one by one. Content that was sitting dormant in libraries suddenly has a second life.

Li Zhao puts commercialization plainly: IP value comes first.

The old hit logic is that a company pours huge effort into one premium piece, then distributes it through the internet so everyone hears it, likes it, sings it, and forms a shared musical memory.

Jay Chou is the example of the early model.

Now the contexts of listening have fragmented. Showering, commuting, cleaning, exercising—people pull up different playlists.

So many songs are produced every day that pushing one premium track to everyone may mean fewer chances of being heard. People look for music by the mood they want right now.

The payment logic will change too. When someone is willing to pay for a song, they may first be buying the premium work and the IP—the person they like, their feelings, their life.

If you like Mao Buyi or Tang Tian, you care about what they’ve been through and what they express.

On the other side, an individual may also pay for the musical atmosphere needed in a particular scenario.

Li Zhao’s judgment: long-term paying in the future is more likely to revolve around the people and IPs they love, rather than an isolated track.

Song Kaifa adds two very concrete incremental scenarios.

One is song translation. They’ve worked with a Hebrew-language singer who sings beautifully but doesn’t speak English; she wants to translate her songs into English and other major languages using her own voice, and is willing to pay 100–200 USD per minute.

Right now voiceover can translate rap and deliver it in her voice; preserving singing performance isn’t there yet, and they’re researching it.

The other is real-time interpreted dubbing. Gamers want their voice to instantly convert into a language their teammates understand while they speak; at events like the ASEAN Expo and the Xi’an Investment and Trade Fair, people want to speak in their own voice for simultaneous interpretation.

On the next visit, the other side should still recognize that district mayor’s or governor’s original voice.

If this path works, smart glasses, games, and toys will all change. The more personalized the voice, the higher the scenario value and pricing power—no need to keep dragging voiceover prices down to a few yuan a minute.

Xia Yong closes from an investor’s perspective: tool companies are judged more by whether they embed into users’ workflows.

Can one person’s need be abstracted into an industry’s need? Do that, and you have industry value.

He mentions a detail: a creator wrote and performed a song about Chengdu in one night.

In the past that might have taken dozens of people. Today they can do it because they had already accumulated IP imagery and digital assets; when expression was needed, those capabilities were there.

That’s the key. Single content ages; workflows, digital assets, and IP relationships stay.

Xia Yong has invested in Ximalaya, Original Force, Sparkly Key, and Shushu Group; his read on content companies is unsentimental.

Startups need cost and efficiency gains to survive, he says. But from an investor’s perspective, that may be a relatively short-lived capability.

If everyone keeps competing on price and marginal productivity, long-term value may not be high.

His harshest judgment is on premium-content companies.

“Companies that do pure premium content may be better suited to platform ecosystems with distribution power, like Kuaishou or ByteDance.”

Content carries the creator’s worldview and imagination, but premium work can’t be infinitely replicated. The audience expecting Ne Zha 3 doesn’t mean you can keep going to Ne Zha 10.

In April this year he attended the China Network Audio-Visual Conference; there were more than twenty projects from around the country, and many institutions looking, including film, game, and AI-film projects.

Many institutions didn’t pull the trigger. To be honest, the certainty of investing in premium-content production companies may be lower than for tools and platforms.

So what can be invested in?

Xia Yong’s answer is consistent: technical iteration is the threshold and the passing grade. What matters more is cognition, emotional expression, and worldview construction.

Entry-level technical roles may lose value, but true experts become more expensive.

Not long ago, a data-annotation company was even hiring PhDs at 30,000–50,000 RMB a month for cognition-related work.

That’s the easiest place to misjudge this wave of AI-content startups.

You think investors want to hear “I can bring costs down.” What they actually want to hear is: Whose workflow have you embedded into? What digital assets have you accumulated? Do you have the ability to keep producing premium work? Can you abstract one person’s need into an industry’s need? When the model keeps getting stronger, what’s left of you?

After technology is democratized, what actually separates players is the things that look slow: taste, judgment, emotion, cultural understanding, IP, and the know-how an industry has built over years.

Near the end of the roundtable, Xia Yong says something plain: even premium-content companies need to keep up with technical iteration, while holding on to their own value judgment, eye, cognition, and feel for beauty.

AI will keep pushing prices down. Songwriting, voiceover, translation, drawing, video generation—all will get cheaper.

But the hardest problem in the content industry has returned to the way it always was.

When a melody can split into 1,800 versions, what is truly scarce is: which version do you want people to remember.

And the more content there is, the scarcer the thing is not the ability to generate, but knowing what is worth generating.

* * *

**Host:** Wu Shenliang, VP of Extraordinary Capital

**Guests:** - Li Zhao, Music Director, Sichuan Observation - Chen Shenhan, Founder & CEO, Infinite Picture Book - Song Kaifa, Co-founder & COO, VMEG (Sudong Technology) - Xia Yong, Partner, Shanghai Xingwang Investment

**Wu Shenliang:** For the next 50 minutes, we’ll discuss how AI defines digital content and the relationship between people and content. Please each take a minute or two to introduce yourself and your industry.

**Li Zhao:** Hello, I’m Li Zhao, music director at Sichuan Observation. Sichuan Observation is Sichuan’s official media, with 175 million users across platforms. Welcome to Sichuan Observation.

**Chen Shenhan:** Hello, I’m Chen Shenhan, founder and CEO of Infinite Picture Book. Everyone knows how Jack Ma went from English teacher to founding Alibaba; I went the other way—out of Alibaba, then spent a time as an English teacher. I read going from English teacher to Alibaba as a humanistic journey toward the center of technology; going from Alibaba back to English teacher is returning to humanism, a call back to first principles.

My team includes members with degrees from the world’s top-ten universities and from Tsinghua and Peking University; we also have a strategic cooperation letter of intent with Alibaba Cloud. Today I’m looking for partner investors and outstanding screenwriters. In one sentence, the company is an experience firm building AI interactive narrative and a scientific decision engine, currently applied in education.

**Wu Shenliang:** Mr. Xia does investing; you can talk after the session. Mr. Song, please.

**Song Kaifa:** Hello, I’m Song Kaifa, co-founder and COO of Sudong Technology, VMEG.AI. We use AI to make voices personalized and take them global. The company is under two years old, 95% of users are overseas, close to 2 million. As Chinese AI short dramas go overseas, we’ve also drawn attention from domestic peers and investors. We want to use AI to bring people’s voices around the world and do multilingual localization.

**Wu Shenliang:** Mr. Xia.

**Xia Yong:** Hello, I’m Xia Yong, partner at Shanghai Xingwang Investment. We focus on cultural consumption and rank roughly top-20 nationally. We’ve invested in Ximalaya, Original Force, Sparkly Key, and Shushu Group; two of them worked on Ne Zha 2 and Eight Immortals. Our institution wants to help cultural confidence advance to a higher stage.

**Wu Shenliang:** Thanks, all four. First question, starting from common ground. In the past, talk of AI mostly meant cost and efficiency: work ten people did, now two can do. But if we understand this wave of AI only as efficiency gains, I think it’s incomplete.

Take music: in the future it may no longer be a fixed three-minute piece, but music generated for each person’s state and scene. Mr. Li, from what you see, is the biggest change to music that it lowers the barrier to creation, or has it already changed the definition of “a good piece of music”—even “a piece of music”?

**Li Zhao:** Before I answer, let me see how many of you here use AI software for music—could you raise your hands? Quite a few. Our observation of the music industry is that, on the surface, AI raises efficiency; underneath, it changes the form.

Traditional music production has long cycles: lyrics, composition, arrangement, recording, post, with more discussion in between and clients giving lots of notes.

With AI software, a lot of content can be generated in one click: one-click lyrics, composition, arrangement, post, even the vocal—you don’t even need the singer to sing. After recording a person’s voice, you can keep generating content, saving enormous time. But that’s just the surface.

In the past music was usually released in fixed forms like singles and albums. Now a good melody, when it enters internet distribution, may already exist in countless versions; you can hear arrangements in different styles. The paths of distribution and generation have both changed.

So from our view, on the surface it’s efficiency; underneath, the form of the song has been changed by AI tools.

**Wu Shenliang:** I agree strongly on form. Chen Shenhan mentioned in an earlier interview with Extraordinary Research that Infinite Picture Book wants children to “live a story,” not just “watch a story.” That also changes how children receive content. Can you talk about what you’re building now?

**Chen Shenhan:** Mr. Li put it well. That AI raises the production efficiency of various art forms is fairly certain; next, each content type has to explore how its form changes.

Take Infinite Picture Book’s interactive picture book: based on large-model capabilities, we string together text-to-text, text-to-image, image-to-image, and text-to-video and image-to-video multimodal capabilities, then put the user inside the content-production process.

When users create themselves, they have more feeling and thinking. Deeply pursued, this can form a new content form.

**Wu Shenliang:** Voice distribution is similar. In the past, Chinese creators wanted to distribute content overseas, but translation and localization were complex. After AI reopened the chain, things got much simpler. Mr. Song, do you think AI mainly changes production cost, or does it widen the reachable boundary for creators?

**Song Kaifa:** Looking at AI voiceover users, I lean toward the latter—it extends users’ capability. The content or IP itself has to make overseas audiences want to watch it, and then you dub it into Arabic or Vietnamese and take it to different regions. AI upgrades this cross-regional distribution capability.

But the AI voiceover market has layers. High-end needs still sometimes can’t be replaced by AI. Today’s going-overseas AI short dramas have relatively few roles—male lead, second male, female lead, second female, plus supporting cast, often under ten. Real film and TV have many roles and deeper emotional performance; AI can’t do that yet.

AI hits low- and mid-tier localization hard, making previously stuck content flow. A Shanghai client told me some Russian animations suit Chinese audiences but couldn’t enter before; now they can be dubbed into Chinese and distributed domestically.

Similarly, “One Belt, One Road” content going to Africa—a Swahili version, or a Kazakh version elsewhere—was hard to do quickly by traditional human means. AI drives the flow of existing content; that’s the industry change I feel.

**Wu Shenliang:** Chinese cultural outreach needs more from Mr. Song’s company. Mr. Xia, from an investor’s view, if a startup says “what used to cost 100 now costs 10,” is that enough for you to invest? What capability matters more when you invest in content companies now?

**Xia Yong:** Startups need cost and efficiency to survive and grow. But from an institution’s view, that may be a relatively short-term capability. If everyone keeps competing on price and marginal productivity, long-term value may not be high.

We care more about the tool’s impact on the industry. For AI content, AI music: can its user network cover enough users, can it truly embed into MCN and creator workflows, and grow everyone’s digital assets. That may have more lasting effect than short-term cost cuts.

**Wu Shenliang:** The key is how deep it embeds into real workflows. Second question: how should AI and humans divide work today? Mr. Li, people who don’t know music theory, with no formal training, can now use AI to make a listenable song. Where does the value of a professional musician live in the future?

**Li Zhao:** First define AI: I see it as a tool, not a product designed to replace people. We hold the right to use AI software. In the past the barrier to entering music was high; people studied music from childhood and had a specialty. Now there are different ways of producing.

We followed a delivery worker who loves music and has no professional training. He picked a niche: chat with people, hear their stories, turn them into songs and post online.

Unexpectedly those songs got big traffic; then more and more people asked him to write songs—government agencies, scenic spots; people asked him to write a song to find a lost daughter or to propose to a girlfriend. He can earn over 100,000 RMB a month; that was beyond what I expected.

Others use AI to cover classics and get traffic; maybe they can’t monetize through copyright, but they may earn through performances.

Tools let different people use their strengths. The delivery worker really understands people; he finds the expressive point in others’ stories, then borrows AI to create.

But professional musicians can’t drop music theory and fundamentals, or they lose taste. If everyone generates with one click, music gets more homogeneous.

To make better work in the future, beyond music theory and understanding of timbre and humanistic expression, what matters most is taste and the ability to choose good songs.

Musicians actually need to learn a broader range: software operation, how to choose the right music with keywords, and judging whether generated content is reasonable.

For example, making a song about Sanxingdui and AI adding a guzheng—you have to consider whether the guzheng matches the era of Sanxingdui. That needs knowledge, taste, and continuous learning.

**Wu Shenliang:** Ordinary people extend their creative ability with AI; professional musicians, with taste and expertise, judge what’s better and how to make it better. Shenhan, you build for children but don’t let children generate everything alone—adults or professionals control the product while children participate in the experience. What’s the difference between AI “doing it for the child” and “helping the child”?

**Chen Shenhan:** Mr. Li’s example shows this is a great era for individual expression. People ask me whether AI will destroy or liberate humanity. I don’t think it depends on AI. The large model provides breadth, width, and a huge content pool; the key is how a person distills their own thing from it, and with what mindset.

Infinite Picture Book now works in education; users are children. Children have rich imagination but limited expression. Our goal is to spark imagination and provide space, while releasing their limited expression.

Through story interaction and plot choices, children form their own feelings and thoughts; in the visualization step we dig out what they want to express, making the work part of their own experience.

The taste, feeling, and logical thinking produced in this process are what can truly integrate into growth. In the past, lecture-style education had adults teaching children with their own experience; without lived experience, the effect is limited.

**Wu Shenliang:** Thanks, Shenhan. AI extends human capability, and users build taste along the way. Mr. Song, translation, voiceover, and voice-cloning model capabilities will keep getting stronger. For a company like yours, where can long-term advantage be built?

**Song Kaifa:** Translation and voiceover need to be discussed separately. Large models have hit the translation industry hard, and traditional translation companies compete fiercely on price. Our core is voiceover.

Big clients often do their own translation and polishing, then hand the script to us to dub documentaries, TV, short dramas, and training/education content.

I personally think long-term, voiceover may not have a technical moat either. As long as a language has enough corpus, large models may eventually do it all.

Right now some low-resource languages, or domestic Tibetan and Uyghur, still lack mature models; but if models cover them, middle companies like ours may thin out. Users can just have the model generate a video speaking a given language with a voice.

Right now there are still thresholds because humans still lead the content. We also deal with regulators and premium-drama clients’ requirements—you can’t hand every voice to AI.

For example, professional voice actors are expensive and time-limited; you can have other real voice actors perform first, then AI correct the shortcomings. If 80–90% of the emotional expression still comes from a real human, it won’t sound like the AI voice that over-enunciates.

Human-AI combination needs people who understand the voiceover industry and the specific content; that’s what we can do well now.

Localization also has thresholds. Korean and Vietnamese have regional differences. Clients ask, which regional accent? If you can’t even tell the target market’s language habits, it’s hard to do. We choose the right local voice based on content and region.

Also, most AI Voice or TTS demos show one person speaking, but real production may be four, five, seven, or eight people in dialogue, even in noisy scenes. Each needs a unique voice and performance.

Single-person voice conversion is one thing; dubbing complex multi-person scenes is another. The latter still needs people and commands different price tiers.

So there’s a short-term moat, but long-term it may shrink. That’s my read.

**Wu Shenliang:** Mr. Xia, as an investor, do you recognize the moats Mr. Song describes? As model capability keeps strengthening, where is a company’s long-term moat?

**Xia Yong:** This reminds me of Ximalaya when I was there—people were at four or five thousand. With these capabilities, you might need a tenth of the people.

In the past, communicating with technical people was hard; now many things people can do themselves. Large models rapidly drove technology democratization.

So at least at this stage, technical iteration is the threshold and the passing grade. What may matter more is cognition, emotional expression, and worldview construction.

Mr. Song’s multi-person complex scenes—there’s a form called audio drama. I personally think even with stronger models, in three to five years it’s hard to fully replace.

When investing in AI content companies, either look at platforms and tools like Ximalaya, or at whether it has the ability to create or stock IP—those may form longer-term competitiveness through technical iteration.

Entry-level technical roles may lose value, but true experts become more expensive. Not long ago a data-annotation company was even hiring PhDs at 30,000–50,000 RMB a month for cognition work.

Technology reaching the passing grade, plus emotional expression and worldview construction, is what can form mid- to long-term competitiveness.

**Wu Shenliang:** Thanks, Mr. Xia. Next, user experience and commercial value. The old content-industry hit logic was to make a product tens of millions of people have seen.

In the AI era, a company may produce lots of content matched to each person’s needs. Shenhan, does Infinite Picture Book want to make picture books most children like, or a bespoke picture book for each child?

**Chen Shenhan:** The two aren’t in conflict; they even reinforce each other. Universal education has run from the industrial era and has its necessity. All children need general knowledge, basic concepts, numerical ability, and a basic feel for beauty—that can’t be replaced.

But personalization can’t be missing either. In the AI era, if you only train children to test and memorize, they may end up competing directly with AI.

Many parents ask me what classes to enroll in and what education to invest in; children are still super competitive at university. The root issue is that the structure and proportion of educational content need to change. Universal education stays, and personalized growth is needed too.

Infinite Picture Book wants to provide personalized content companionship, helping children know themselves, discover talent at a young age, form a healthy personality, and face an uncertain era.

At the same time, popular content is needed to carry traditional culture and shape the company’s IP core, so all children see a shared world.

Our goal isn’t to make every child completely different; it’s that each child, growing up, can find their own place in this shared world.

**Wu Shenliang:** Music is similar. In the future, we’ll need musicians like Jay Chou providing classics everyone loves, alongside bespoke music tied to personal preference and upbringing.

**Li Zhao:** From commercialization, we think IP value comes first. One-click AI content tends to be homogeneous; how do you keep users?

The old hit logic is that a company pours huge effort into one premium piece, distributes it online, and everyone hears it, likes it, sings it, forming a shared musical memory. Jay Chou is the early-model example.

Now listening contexts have changed. Showering, commuting, cleaning, exercising—people pull up different playlists.

So many songs are produced daily that pushing one premium track to everyone may mean fewer chances of being heard. People look for music by the mood they want right now.

The traditional path is premium production, promotion, clicks; now a melody at promotion may already have spawned 1,800 versions entering different fields and scenes—some for children, some for other vertical scenes.

Music consumption moves from mass to vertical, and may further move to personal customization.

Payment logic also changes. When I’m willing to pay for a song, I may first buy the premium work and the IP—the person I like, their feelings and life.

If you like Mao Buyi or Tang Tian, you care about what they’ve been through and express. On the other side, an individual may pay for the musical atmosphere needed in a scene.

I think long-term paying is more likely to revolve around the people and IPs they love, rather than an isolated track. That’s my view.

**Wu Shenliang:** AI makes music form infinitely variable, but what moves people is still the emotion and story behind the song. Mr. Song, from what you see, what opportunities or challenges come with the spread of voice technology?

**Song Kaifa:** Building on Mr. Li, I’ll share two directions. First, song translation.

We’ve worked with a Hebrew singer who sings beautifully but doesn’t speak English; she wants to translate songs into English and other major languages using her own voice, and is willing to pay 100–200 USD per minute, asking whether we can do it.

Right now we do voiceover; if it’s rap, we can translate and deliver in her voice; preserving singing isn’t there yet, we’re researching.

We also have a user in Nepal using our tool who wants to take their songs to other regions. Musicians are willing to pay for this; if the tech breaks through, I think it’s an opportunity market.

Second, real-time interpreted dubbing. Gamers want their voice to instantly convert into a language teammates understand, and to understand teammates back.

At events like the ASEAN Expo and the Xi’an Fair, people want to speak in their own voice for simultaneous interpretation, not just hear a stranger interpreter’s voice. On the next visit, the other side still recognizes that district mayor’s or governor’s original voice.

If we get there, smart glasses abroad both help me understand others and deliver what I say in my own voice.

Similar ability can enter games and toys, giving smart toys a personalized voice. We do voiceover and have been moving toward personalization.

As personalization rises, scenarios and pricing power rise too; no need to keep dragging voiceover prices to a few yuan a minute. In the future we want to combine AI voiceover with AI songs for these incremental scenarios.

**Wu Shenliang:** Very on point. The audio ecosystem still has lots of room; iteration deepens the scenarios. Mr. Xia worked at Ximalaya. Do you think future content-company business models will change? For example, as Mr. Li says, users may pay for IP, people, and stories, not a single track.

**Xia Yong:** What everyone said is good. Today’s AI digital-content companies may emphasize prompts, API calls, databases; those are competitiveness at a certain stage. But tool companies are better judged by whether they enter users’ workflows.

Can one person’s need be abstracted into an industry’s need? Do that, and you have industry value.

Film and music works carry IP. In the past the IP might be a person, or a song; in the future it may be richer.

As digital assets accumulate, we see a creator writing and performing a song about Chengdu in one night. In the past that might have taken dozens; today they can because they’d accumulated IP imagery and digital assets; when expression was needed, those capabilities were there.

AI music and audio may be relatively niche, but in digital content they can be a value amplifier. Dialect and regional accents make characters more vivid; the experience of one work in uniform Mandarin differs from versions with regional accents.

The more diverse AI digital content is, the more value it may create.

**Wu Shenliang:** I also agree on the possibility dialect brings. An audience member asks: from an institution’s view, if a pure premium-content production company wants to raise funding, what advice, what conditions?

**Xia Yong:** Personally, I think pure premium-content companies may be better suited to platform ecosystems with distribution power like Kuaishou or ByteDance.

Content carries the creator’s worldview and imagination, but premium work can’t be infinitely replicated. The audience expecting Ne Zha 3 doesn’t mean you can keep going to Ne Zha 10. So premium-content companies need a distribution ecosystem where they can grow.

Also, you must embrace technology change. I know two teachers surnamed Leng, originally in traditional industry training; when AI came they quickly transformed and became prominent in this field.

Premium content must also keep up with technical iteration, while holding on to your own value judgment, eye, cognition, and feel for beauty.

In April this year I attended the China Network Audio-Visual Conference; there were more than twenty projects from around the country, and many institutions looking, including film, game, and AI-film projects.

Many institutions didn’t pull the trigger. Honestly, the certainty of investing in premium-content production companies may be lower than for tools and platforms.

For example, when we invested in Original Force, two things mattered: first, long-accumulated ability to keep producing premium work; second, through such production houses, earlier exposure to projects that might break out.

But projects go through multi-layer screening, and content companies need ecosystem support to grow. In the end, you have to hold your own value judgment.

_Roundtable recorded at the Extraordinary Awards Chengdu AI Entertainment Conference, September 2026. Transcript adapted by Extraordinary Research (非凡产研)._

[![cover](https://substackcdn.com/image/fetch/$s_!DqkC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98074fd-48ff-4e09-b0d0-c4ba5c1fe194_2730x1536.jpeg)](https://substackcdn.com/image/fetch/$s_!DqkC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98074fd-48ff-4e09-b0d0-c4ba5c1fe194_2730x1536.jpeg)

---

Original publication: https://uniqueresearch.substack.com/p/after-ai-pushes-content-to-10-yuan
On-site reading page: https://ffcap.cn/en/research/after-ai-pushes-content-to-10-yuan
