Original · Unique Research · 2025-11-28
Editorial note: This is the complete historical article and five-round panel conversation published in November 2025. Adoption figures, performance scores, musical-superiority claims, generation times, upcoming releases and the planned 40% price cut are attributed to the source and speakers, not independently verified current facts. Claims that copyright becomes worthless and that a melody is original merely because it has not appeared before express the panel's economic or creative opinions, not a statement of copyright law. The source attributes the claim to have coined MaaS to Yang Yue. The moderator refers to Sudu Technology and says it has reached 99.9 points; Song Kaifa, introduced as VMEG's co-founder, responds that they are still working toward that score. Both the source's inconsistent company reference and the correction are preserved, without assuming the two company names are equivalent. Relative dates and predictions refer to the historical conversation.
If one day 99% of the world's music is not written by humans, will you still play a song on repeat?
When short-form dramas use AI soundtracks, films use AI dubbing, and every piece of background music (BGM) behind gyms, yoga studios, and livestream rooms is generated by algorithms in real time, what will become of music, the most emotional of businesses?
At a roundtable during the 2025 Unique Bloom Beijing, we put questions like these to four people working on the front lines:
The moderator was Fan Zhihui, founder of Music Business China and Entertainment Business China, who has spent years closely observing the industry;
Song Kaifa, co-founder and COO of VMEG, works in global AI dubbing and spends his days refining voices with personality alongside Indian production companies and overseas television stations;
Cheng Zhaoyu, Music Director at multimodal foundation-model company MiniMax, is thinking about one question: when text, video, speech, and music are all entrusted to the same model, what entirely new forms of creation will emerge?
Guo Rui, Head of Growth Marketing at Mureka, which provides music models and APIs to musicians and developers, is experiencing the real pressures of commercialization and cost curves firsthand;
And Yang Yue, both a musician and an AI entrepreneur, is co-founder of 43AI Technology Group and CEO of 43Music. He advanced the controversial view that 99% of the music humanity ought to hear has not yet been made.
This article is not a simple verbatim transcript. It reconnects the views, disagreements, and moments of insight from that occasion to ask: when AI starts writing songs, how many of the music industry's rules will remain the same?
I. From Announcer-Style Delivery to a Human Feel: The Dubbing Industry Is Being Rewritten
The discussion began with the most practical business: AI dubbing.
Song Kaifa said that VMEG specializes in dubbing content for the world: short-form dramas, films, animation, advertising, and documentaries. It currently has more than 1 million users, including Indian production companies, overseas television stations, and film companies.
Two words appear most frequently in feedback from users around the world:
Natural and Real.
Traditional AI dubbing, especially models from many large companies, generally sounds like a broadcast announcer: standard, clear, flawless, and unsurprising. Most voices are simply those of adult men or women, with emotions delivered in a flat, straightforward way.
The real world needs far more complex voices. Children, elderly people, dialects, accents, and even a highly specific voice such as that of an elderly Tibetan woman are all scarce resources.
What VMEG is doing can be summed up in one sentence: filling out the long tail of voices across more than 170 languages.
It is not enough merely to speak; the voice also has to sustain long-form speech.
Many AI dubbing products have no trouble reading a commercial lasting a few dozen seconds, but their weaknesses show in long videos: the timbre becomes inconsistent, the emotion suddenly shifts, or the character's persona falls apart.
Most of Song Kaifa's customers, however, test the product with two-hour films. Those films also contain songs, emotional climaxes, and extensive dialogue, so the slightest mistake breaks the illusion.
That is why they have concentrated their moat in two areas: timbre and emotion.
When a virtual voice can maintain a stable persona for hours and transition naturally through crying, laughter, anger, and silence, dubbing is no longer merely comprehensible; it becomes something audiences can truly engage with.
Imagine the future: an Indian film can be dubbed automatically into North Indian languages, South Indian languages, and multiple other languages;
For the songs, the producer can choose between a dubbed cover and retaining the original vocals;
All of this is embedded in the production workflow, leaving the director to focus only on the result.
Then you understand why Song Kaifa stresses that the true moat is not knowing how to speak, but understanding film.
II. If Music Loses Its Scarcity, Will the Logic of Copyright Collapse?
If dubbing is repairing efficiency gaps in the existing industry chain, what 43Music is doing comes closer to rebuilding the logic itself.
Yang Yue's view is direct:
Over the past year and a half, people's view of AI music has changed from an efficient tool to a professional assistant and then to today's reality: leading models can already outperform the overwhelming majority of professional musicians.
He left a little room, saying the remaining 1% was there to preserve some human dignity, but his meaning was clear:
At the level of creation, AI is no longer a toy. It is genuine productive capacity.
That raises a question: what is the traditional music industry's foundation of value?
The answer is simple: scarcity.
Writing songs takes time, experience, inspiration, and opportunity cost.
That is why the copyright to a hit single can keep generating income for its creator.
But AI has shattered all of that.
When models can produce high-quality music at scale in a short time, writing a song is no longer scarce, and copyright naturally cannot command the same price.
You might write ten songs in a year, have one become a hit, and live on the copyright for ten years;
AI can help you generate one hundred songs a day, each of them fairly good, but none with a reason it alone must exist.
Yang Yue therefore proposed a new concept: MaaS—Music as a Service.
Under this logic, music is no longer sold as individual works. It becomes a service capability tied to a particular setting:
A yoga studio receives a complete set of music aligned with breathing rhythms and the rhythms of muscular relaxation;
An exercise setting receives music whose tempo adjusts dynamically according to heart rate and performance feedback;
Sleep and meditation receive an audio prescription that reduces tension and helps people fall asleep;
Study and concentration receive background sound that maintains a rhythm without demanding too much attention.
Music is no longer a matter of listening to whatever is available; whatever you need is what I generate.
More interestingly, they do not even impose a strict anti-piracy strategy.
The reason is practical:
Making a song takes only fifteen seconds,
While stealing one still requires four minutes to transcribe, import, and clean it.
In an era when production costs approach zero, preventing theft itself becomes uneconomical.
What truly has value is no longer that one song, but whether you can continuously provide genuinely useful audio solutions for a specific setting.
III. A Multimodal Company: Why Go to the Trouble of Making Music?
If you are a foundation-model company already competing in text, images, and video, why move into music, a field that seems niche and difficult to evaluate?
MiniMax's answer is resonance across modalities.
Cheng Zhaoyu noted that in a multimodal architecture, text, speech, video, and music are not four parallel lines. In many places, they can share underlying experience and training frameworks.
More importantly, as models become more capable, the boundaries are blurring.
It is increasingly difficult to separate a model designed specifically for music from one designed for video, because an audiovisual work inherently includes images, dialogue, sound effects, and a soundtrack at the same time.
At the product level, they are now working on one-click creation:
With an idea, a passage of text, or a simple prompt, you can generate a complete song.
For most listeners, it is already difficult to distinguish by ear between an AI-generated pop song and one made by humans.
The real challenge is not whether the timbre sounds convincing, but two things: aesthetic quality and originality.
Aesthetic quality can be learned from large amounts of data. Put bluntly: do not make it too unpleasant, too jarring, or too far removed from basic musical taste.
Originality, however, is troublesome.
In theory, as long as a melody you write has never appeared anywhere in the world, it is original.
In the real world, however, we all know the question: why do so many songs sound alike?
Because they merely swap a few spices in the recipe; fundamentally, they are still prepackaged dishes.
MiniMax is using an evaluation system to answer two questions:
Does a song conform to broadly shared aesthetic standards—at least enough that it is not unpleasant to hear?
And can it simultaneously give the creator a sense that "this is mine," rather than sounding like a template assembled from a collection of formulas?
When multimodal capabilities are layered together, music is no longer merely BGM added afterward.
It can be designed together with images and text, becoming part of the overall content experience.
IV. From Model Iteration to a 40% Price Cut: A Real Turning Point for B2B Business
It is easy to get excited when discussing technology, but eventually all that excitement comes down to one question: how hard is it to make money?
Mureka's answer is highly pragmatic. The company develops music products for creators while also providing developers with APIs and model services.
Guo Rui revealed that its forthcoming updates to the O Series and V Series models will be more nuanced overall:
They will not only understand broad styles such as jazz or rock, but also subtle emotions such as a lazy afternoon with a faintly broken heart that has not completely fallen apart.
On the B2B side, the models will be more stable, faster, and easier to control. Then came a crucial statement: the new generation of models will be priced 40% lower overall.
Why?
Because they have found that more and more B2B companies are integrating music models:
Companies working on AI video, traditional video advertising, games, virtual idols, and more.
These customers' demands are entirely different from those of independent musicians. They care more about:
Can it produce output reliably?
Is the latency low enough?
Can the pricing structure allow the entire ecosystem to make money?
If a model provider places itself at the top of the ecosystem but keeps prices rigidly high, leaving the circle of application developers below it unable to make money, that ecosystem will quickly wither.
The price cut is therefore not charity but a sober judgment:
Once technology crosses a certain threshold, sharing gains appropriately can instead expand the entire market.
V. The Memory of a Voice: Dubbing Does Not Replace Emotion, but Amplifies It
Returning to voice itself, it is more than a carrier of information. It is also emotional memory.
Recently, dubbing issues surrounding Zootopia 2 have sparked intense discussion in China: audiences do not dislike celebrities; they simply do not want to give up their memories of classic characters' voices.
While providing dubbing overseas, Song Kaifa has seen more specific needs: some musicians want to sing in languages they do not speak at all, such as Hebrew or Lithuanian, while preserving their own voice; celebrities and IP owners want their original vocal personae to speak languages around the world; students want to hear familiar teachers deliver courses in different languages, because even if the audio is AI-generated, the learning experience feels more approachable.
These needs all point to one fact: voice is an index to memory.
When you hear a familiar voice, your brain automatically summons the emotions associated with it.
This is also why tolerance for AI dubbing differs enormously between regions.
In some Latin American countries, viewers are highly receptive to AI-dubbed short-form dramas; in North America, however, many users abandon a series as soon as they sense that the voice may not be human.
To keep AI voices from breaking the illusion, VMEG is tackling capabilities that may seem like small details but are essential to the experience: an AI crying voice that truly trembles and cracks; simultaneous speech that gives several overlapping voices spatial depth; and, in the future, conference interpretation in which a leader speaks Chinese onstage while each person in the audience hears that leader's own voice speaking English, Japanese, or Spanish.
These involve technical challenges, but the essence remains the same: AI is not here to eliminate people; it amplifies human emotion so it can reach a larger world.
VI. The Tools of Creation Have Changed, but What It Means to Be a Musician Has Not
Should ordinary music professionals be anxious?
Yang Yue gave an unsentimental answer: the claim that AI means everyone can become a musician is nonsense.
Music has always been the work of a very small minority.
The old barriers were technique, music theory, instrumental skill, and production ability; the new barrier is something else.
Are you willing to keep learning and continually upgrade how you work?
AI's arrival is simply raising the industry's average score:
Someone once capable only of making an 80-point work now has a chance to use tools to make a 95-point work;
Someone who could once reach 100 points now has a chance to reach 120.
One result of this process is that a great deal of bad music will be eliminated.
But that does not mean everyone has lost their opportunity.
On the contrary, people with a genuine desire to express themselves and a distinctive perspective will have a greater chance to break through.
Cheng Zhaoyu said that individual creators of the future may all become super individuals: one person will be a copywriter, director, composer, and producer at once.
AI is not writing the song for you; it is helping you turn the flash in your mind into a complete work with the least possible friction.
As long as you have something to say and emotions to share, you will always have an opportunity.
VII. The Underlying Logic of Growth: Return to People Before Talking About Growth Hacking
When the discussion turns to commercialization, growth naturally cannot be avoided.
Mureka's view, however, is that growth will inevitably go astray if it focuses only on data and hacking techniques while losing touch with creators and users themselves.
From its perspective, there are two main groups of users. The first consists of creators and musicians, whose true need is expression.
They may not care whether a song charts. They care more about whether there is a tool convenient enough to help turn the images in their minds into something that can be heard.
The other group consists of B2B customers and developers. Their need is not romance, but control and stability.
Can they make stable calls through an API?
Can output quality remain consistent every time?
Can parameters be tuned precisely for a business setting?
One phrase mentioned repeatedly here was Domain Expertise.
AI will not automatically understand what a yoga instructor, fitness coach, or advertising company actually needs.
It takes people who understand music, products, and industry settings at the same time to translate technology into the language of business.
Growth is not a set of formulas for acquisition, conversion, and retention; it is whether you can keep helping a group of people solve real problems.
VIII. Advice for Entrepreneurs Entering AI Audio: Build Moats, Be Inimitable, and Let Go of Yourself
Near the end, Fan Zhihui asked the question everyone had been thinking about: if you were entering AI audio today and could each offer one piece of advice, what would you say?
The four guests offered four perspectives.
Song Kaifa's advice was: understand the setting thoroughly before considering the moat.
The AI voice market is enormous. From places such as Russia, where even basic translation tools are underdeveloped, to short-form drama, film, and education, gaps exist everywhere.
But if all you make is a general-purpose model that anyone can build, you will eventually fall into homogenized competition.
They have chosen a personalized-agent path:
When dubbing for a person of a particular age, accent, and personality, the system can find a highly precise match.
The deeper your understanding of a particular setting, the harder it is for users to leave you.
Cheng Zhaoyu's advice was: follow a path that cannot be copied.
Do not compete with large companies on production polish; they will always have more professional teams.
What you should compete on is your ability to continuously reveal your own inner core.
He gave an interesting example:
On Bilibili, there is an uploader described as a remix creator of Douyin's Sister Yu. The creator essentially pairs fragments of everyday life in Northeast China with hypnotically catchy melodies.
Is it professional? Perhaps not from an academic perspective, but that lived-in texture mixed with a little amateurism cannot be imitated by others.
Individual entrepreneurs should move freely among video, text, and music to build their own distinctive expressive style,
Rather than dreaming that one viral AI song will make them famous overnight.
Guo Rui's advice was to place yourself in the role of a tool provider: what creators truly want is to be seen by the world,
So identify the obstacles they encounter at each step and help remove them one by one.
Yang Yue, meanwhile, offered completely different advice to two groups:
For musicians preparing to use AI, he advised learning to let go of the melodies they already have in mind.
Do not obsess over making AI perfect the short passage you hummed. Instead, try surrendering control and handing the act of generation over to AI,
And you may see a world far broader than what was in your own head.
To traditional musicians who firmly reject AI, he instead said: then please persist all the way.
The future will be full of human-machine co-creation, and music written by purely carbon-based humans will become an exceptionally scarce luxury.
If you can prove that an album involved no AI at all, you may instead be able to sell it at a higher price.
Those who embrace AI have an opportunity;
Those who reject AI but pursue that choice to the utmost also have an opportunity.
The people in real danger are those who neither embrace nor persist, but simply wait passively.
Epilogue: Human Imperfection May Be the Last Copyright
At the end of the discussion that day, Fan Zhihui said something that sounded like both a conclusion and a reminder:
This era does not require everyone to embrace AI.
What it truly requires is that, whichever path you choose, you follow it to its fullest extent.
AI music and audio technology will keep advancing: voices will become more realistic, emotion more nuanced, generation faster, and prices lower.
But there are some things it will probably never be able to replace.
The breath mixed into an imperfect voice,
The extra press of the delete key when a songwriter hesitates,
The few notes missed from nerves during a performance.
Human creation may shift from the lead role to become one half of a performance shared with AI;
But it is precisely this imperfection, instability, and unpredictability
That constitutes our final scarcity.
AI helps us send voices farther and make more music;
What we must do is decide what kind of voice we truly want to leave to the world in an age of infinite replication.
More Details from the Conversation
Opening Introductions
Fan Zhihui: Hello, everyone. It is a great honor to discuss this topic here with several industry entrepreneurs and experts. I came with a great deal of curiosity and many questions. First, I would like each guest to introduce themselves in turn.
Song Kaifa: Thank you to the Unique Bloom for the invitation. We are a company focused on global expansion, primarily providing AI dubbing, and we now have more than 1 million users worldwide. Indian production companies and many overseas film companies and television stations use our AI tools for dubbing and voice recreation, with the goal of giving voices a more human feel. Today I will share some of our work in global AI dubbing. I am Song Kaifa. Thank you.
Cheng Zhaoyu: Hello, everyone. I am Cheng Zhaoyu. We are MiniMax, a foundation-model company covering video, text, speech, and music. Today I will mainly discuss our work in music.
Guo Rui: Hello, everyone. I am Guo Rui from Mureka. We have music models, music applications, and music API services. It is great to meet everyone today.
Yang Yue: Hello, friends of Unique Research. I am Yang Yue. 43AI Technology Group is an integrated group combining AI R&D, implementation, training, and investment. Within the group, I am mainly responsible for music generation and product implementation. We now have 3 AI applications listed across the major app marketplaces, and I am very happy to discuss AI music with everyone today.
Round One: Technical Barriers and Product Differentiation
Fan Zhihui: I will begin with a question for Mr. Song. In preparing for this discussion, I saw a report saying that, in your view, AI dubbing does not have an 80-point standard—only 99.9 points. From the standpoint of technical practice, what is the greatest technical barrier in AI dubbing? If Sudu Technology has already reached 99.9 points, how do you avoid being overtaken by other companies?
Song Kaifa (VMEG): Thank you, moderator. In fact, in dubbing we are working toward 99.9 points. In the feedback we receive from users around the world, the two most frequent words are "Natural" and "Real."
We develop AI to give voices personality. Today, AI dubbing from many large Chinese companies or overseas companies mostly has an announcer-style delivery, and most voices are those of adult men or women. Globally, however, children's voices, elderly voices, and even specific voices such as those of elderly Tibetans are extremely scarce. We match a wide range of timbres across more than 170 languages worldwide, allowing us to meet a user's need for a particular kind of voice effectively.
As for our moat, to avoid homogenized competition we have continued to go deep in two areas: timbre and emotion. You can tell whether a company's dubbing is good by testing it for 20 to 40 minutes. Many dubbing companies encounter timbre inconsistencies over long durations, whereas our large customers typically produce films that are 2 hours long.
For example, Indian films are made in North Indian and South Indian languages, and producers use our tools to translate and dub them. Indian films often include songs, and we can automatically handle either cover versions or preservation of the original vocals. We have done extensive work at the foundational audio layer while also going deep into workflows for film, television, animation, advertising, and documentaries to help users improve ROI (return on investment).
Fan Zhihui: Thank you, Mr. Song. Next, a question for Mr. Cheng of MiniMax. MiniMax is a typical multimodal company. Compared with a purely vertical AI music company, what advantages come from making AI music within a multimodal company? Do you encounter internal competition for resources?
Cheng Zhaoyu: Multimodality certainly requires each model to consume some training resources. Cooperation within our company is quite good. There are real advantages: for example, text, video, and music can share many training frameworks, bodies of knowledge, and experiences.
Especially as many models iterate, their boundaries are gradually weakening. Although the way we process data and define evaluation standards differs, overall it feels as though, despite doing everything, we are building one large model.
Evaluation itself is a subtle role because music is highly subjective. We previously discussed the problem of prepackaged dishes. In theory, there is no such thing in music: a new melody you write is original as long as it has never appeared before. Why, then, do people feel that so many songs sound the same? Because they lack a core expression that moves people.
Our evaluation addresses two questions:
General aesthetics: This is a standard that appears complex but is the easiest to define.
Originality: Make creators feel that a work is unique, not a collage of prepackaged pieces but a burst of inspiration.
Fan Zhihui: One major role of AI music models is allowing nonprofessionals to capture an inspiration and turn it into a complete work.
I would like to ask Guo Rui of Mureka: I hear you will release a completely new foundation model next week. What are the core upgrades? Can you give us a preview?
Guo Rui: I am happy to share a little. Both our O Series and V Series will receive updates.
In more intuitive terms, the new models are more nuanced. We express our understanding of things at a finer level of granularity. For example, we have strengthened semantic understanding, so the models comprehend not only style but also emotions, such as a lazy afternoon. Audio-quality processing and arranging styles will also become richer.
The V Series serves developers on the B2B side; in commercial terms, it will deliver higher performance, greater stability, and faster speed. Another major change is a 40% price cut, leaving more room for commercial users to earn a profit.
Fan Zhihui: What is the reasoning behind the 40% price cut?
Guo Rui: We found that technological iteration has crossed a threshold. There are now more B2B users in the broader music field, including companies working on AI video, conventional video, and games. Their needs differ from those of pure creators: they require stability, high performance, and enough profit headroom to promote ecosystem development. We hope to make the ecosystem more prosperous.
Round Two: The Nature of AI Music and Changes in Business Models
Fan Zhihui: A question for Yang Yue of 43Music. You understand both technology and the music industry. How do you see the nature of AI creation? What disruptions will AI bring to creation, copyright, management, and distribution in the music industry?
Yang Yue: Those are two questions.
First, the nature of AI music creation. People's understanding of it has changed over the past year and a half. Initially it was seen as an efficiency tool, then as a professional assistant. Today, however, leading models can already outperform 99% of professional musicians. I say there is still 1% they have not surpassed to preserve a hope—a light of humanity.
AI will not stop. It should not merely copy humanity's musical heritage from the past. I strongly agree with one statement: 99% of the music humanity should hear has not yet been made. AI should break through humanity's existing creative mechanisms and styles and create music that did not previously exist in this world. That is where AI is most powerful.
Second, copyright and distribution. The impact will be disruptive. The value of traditional music comes from scarcity—the costs of time, opportunity, and intellect. AI has completely shattered all three.
When AI can rapidly produce vast quantities of high-quality music, scarcity is crushed and copyright becomes worthless. In the past, you wrote ten songs a year and lived on the copyright from the one that became a hit; now you make one hundred songs a day, and none of them has value.
We must therefore find a new anchor of value. I coined the term MaaS (Music as a Service). We do not sell songs or copyrights; we provide setting-specific audio services, such as an electronic-music Party, therapy, sleep, and study. We do not even protect the music from theft, because making one song takes only 15 seconds while stealing it requires 4 minutes of transcription. There is no reason even to steal it.
What I can say with certainty is that, in an era of mass AI production, the traditional logic of copyright is no longer viable.
Fan Zhihui: If the output of 99% of musicians no longer has value, where can they go?
Yang Yue: I have never believed that AI is here to replace musicians. The idea that AI allows everyone to become a musician is nonsense. Music has always been created by a very small minority to serve the overwhelming majority.
AI is a tremendous benefit for traditional musicians. It will raise the floor of human music and eliminate bad music. Someone who previously made 80-point music can now easily make 95-point work; someone who previously reached 100 can now reach 120. But the work will still be made by musicians with exceptional learning ability, not by just anyone.
Round Three: AI Voice, Dubbing, and Emotional Connection
Fan Zhihui: We just mentioned that many arrangers and dubbing professionals may face difficulties. A few days ago, Zootopia 2 faced a dubbing controversy in which celebrities took work from professional voice actors. Mr. Song, what impact will your tools have on dubbing professionals?
Song Kaifa: Audio receives less discussion than visual media in AI, but the human ear is highly sensitive to AI and can easily tell whether a voice feels human.
Overseas, we encounter many needs:
Cross-language singing: Musicians want to sing foreign-language songs in their own voices—for example, in Hebrew or Lithuanian—and they are willing to pay a high price for it.
Preserving original IP voices: Celebrities or IP owners want their own voices to speak languages around the world because voices are memorable.
Education: Students want to hear teachers they know speaking English or Chinese. Even if AI generates the audio, the learning outcome is better. Returning to the Zootopia 2 example, audiences remember the voices of classic characters and may reject a replacement, even if that replacement is a celebrity. The same applies to exporting short-form dramas: Spanish-speaking markets have a high tolerance for AI dubbing, while North American viewers abandon a series when they hear AI. We are therefore tackling technical challenges such as AI crying voices and simultaneous speech, where multiple overlapping voices need spatial depth. In future simultaneous interpretation, a leader could speak onstage while audience members hear that same leader's own voice speaking different languages. This is also an enormous use case.
Round Four: Workflow Integration and Growth Strategy
Fan Zhihui: A question for Mr. Cheng of MiniMax. As a creator and product leader, how do you view integrating AI tools into workflows? How do you balance efficiency and personalization?
Cheng Zhaoyu: Our product currently works with one click: as long as you have inspiration, I will give you a song. In pop music, ordinary listeners already find it difficult to distinguish between AI-made and human-made work.
AI has long existed as an assistive tool—Logic and Ableton, for example, convert physical sounds into rhythm. Today's AI differs in that it represents an entirely new creative approach. It does more than reduce costs and improve efficiency; it turns individuals into super individuals. As musicians connect more deeply with technology, both efficiency and richness will improve.
Fan Zhihui: Guo Rui of Mureka, how do you set growth strategies in global markets? Which monetizes more efficiently, the consumer or B2B side?
Guo Rui: Growth cannot focus only on the numbers in the growth-hacking playbook. At its core, it must return to people and customers' needs.
Musicians/creators: Their need is expression. We need to provide a more convenient way to help them express themselves to the world.
B2B/developers: Their need is control and stability. Domain Expertise is extremely important here. AI cannot replace humans. Only by truly understanding music and users' settings can you achieve good growth.
Fan Zhihui: Mr. Yang Yue, what is the logic behind 43Music's choice of different use-case entry points?
Yang Yue: From last October to now, we have gone through model iteration and explored the boundaries.
First, exploring the limits of AI-generated music.
Second, studying vertical categories. We spent one-third of our time on research into psychology and brain science. For yoga music, for example, we consulted many yoga instructors and found that they too had previously selected music passively. Now we can proactively create music suited to the setting.
For exercise and fitness music, we built an agent to calculate the relationship among physical data, athletic performance, and music.
This reverses the market: previously, people listened to whatever existed; now, they create whatever they need.
Why do we not try to push singles up the charts? Because this is the worst era for music: traffic has been captured by platforms, and promotion costs are extremely high. Commercially, it makes no sense to take AI music produced at zero cost down the high-cost, high-risk path of traditional promotion. We were therefore forced toward scenario-based services, or MaaS.
Round Five: Advice for Entrepreneurs
Fan Zhihui: Finally, please give one piece of advice to entrepreneurs who want to enter AI audio.
Song Kaifa: The AI voice field is vast. For example, even basic translation tools are scarce in the Russian market, so there are many opportunities.
I recommend building moats. Our current moat is personalized translation and dubbing agents. If a director wants content for people of a certain age and with a certain accent, for example, we can match the voice precisely. The deeper your understanding of the setting, the stronger your moat in retaining users.
Cheng Zhaoyu: I recommend taking a path that cannot be copied. Do not compete with viral songs on NetEase Cloud Music or Douyin in terms of professionalism. Compete on your ability to continuously reproduce the qualities at your core.
For example, there is a Bilibili creator who remixes content from Douyin's Sister Yu, pairing everyday life in Northeast China with earworm melodies. People listen because a little amateurism is mixed into the professionalism, and that lived aesthetic is unique. Individual entrepreneurs should combine tools for video, text, and music to create distinctive value rather than chase a one-hit wonder.
Guo Rui: We look not only at technology, but at the point where art and technology meet. Creators' ultimate need is to be seen by the world. As a tool provider, we identify the obstacles creators face and provide solutions that help them express themselves.
Yang Yue: I have two pieces of advice:
For musicians preparing to use AI: learn to let go. Let go of the existing melodies in your mind. Do not try to make AI refine the passage you hummed; the inspiration AI generates is far better than that small idea. Put your ego aside and let AI create completely.
For traditional musicians who reject AI: persist to the end. If you hate AI, do not use it. The future will be full of human-machine co-creation, and music created by purely carbon-based humans will become extraordinarily scarce. If you can prove your music was made entirely by humans, it should command the highest price. Those who embrace AI and those who reject it both have opportunities.
Fan Zhihui: That feels revelatory. This era does not require everyone to embrace AI, but as long as there is enough originality and scarcity, everyone can find a place. Human imperfection may be the final form of scarcity. Thank you to all four guests for your excellent insights!