跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

100 Tracks in One Second: How Can Chinese AI Entrepreneurs Scale the Mountain Called Suno?

Original · Unique Research · 2026-04-12

Editor’s note: This full-text translation preserves the original report, its repeated introductory passage, and the complete six-topic panel with four guests and one moderator. Performance, user counts, market growth, copyright, voice-rights, product-release, acquisition and future-capability statements are attributed to the original author or named speakers; they are not independently verified findings or legal advice. The headline and narrative use 100 tracks per second, while the moderator later says 100 tracks per minute; these are preserved as distinct source formulations, not a verified benchmark. The source’s claim that Google acquired Riffusion, its description of OpenClaw as a large language model, and its model/version labels remain unverified and are not silently substituted with other entities. “Southern Indian language” and “Xinjiang language” are the source’s imprecise labels; no particular language is inferred. Company and personal names are transliterated where an official English form has not yet been verified. “PTSD” and emotional “breakdown” reproduce the author’s rhetoric, not clinical assessments. The original’s awkward “not going around those mountains” wording is retained in the copyright discussion rather than treated as a clear legal proposition. The source date is April 12, 2026; references to April teacher training do not establish the event date or whether that training had already occurred.

Unique Awards

100 Tracks in One Second: How Can Chinese AI Entrepreneurs Scale the Mountain Called Suno?

Chinese Players’ Strategy in Plain Sight

"

When AI’s capacity to generate becomes overwhelmingly abundant and every barrier to entry is flattened,

it instead brings out humanity’s two most valuable trump cards: an exceptionally discerning pair of ears, and the uniquely human emotional breakdown—even one lasting just 0.1 seconds.

Write a song in a minute, complete with a singing voice that could pass for a real person’s. After Suno burst onto the scene, the entire music and audio-production world went through a huge bout of PTSD.

Everyone was asking: if typing a few lines of prompts lets AI generate such polished, catchy pop hits, how are musicians who spent ten years painstakingly learning an instrument and memorizing countless music-theory concepts supposed to earn a living? And that is before considering recent visual models such as Sora and Kling, which now handle simultaneous audio-and-video generation as well.

It sounds like a script for wiping out everyone who works in audio.

But at the AI-audio trends roundtable during Unique Awards · Hangzhou AI WEEK, the picture described by several industry insiders was completely different from the public’s panic.

When AI’s capacity to generate becomes overwhelmingly abundant and every barrier to entry is flattened, it instead brings out humanity’s two most valuable trump cards: an exceptionally discerning pair of ears, and the uniquely human emotional breakdown—even one lasting just 0.1 seconds.

Not Taking Suno Head-On: Chinese Players’ Strategy in Plain Sight

In any discussion of AI music, Suno is a mountain you cannot ignore.

As CTO of Yinchao’s developer, Ziyou Liangji, Jiang Tao is among those most frequently measured against it. But his answer is straightforward: “We are not actually in a competitive relationship with Suno. There is no need for everyone to fight head-on in the same lane.”

Suno is increasingly becoming a complex professional tool, seeking to dominate professional musicians’ workflows. But in Jiang Tao’s view, the general public, film and television, short-form dramas, and games remain a vast underserved market. Before competing with giants on compute, he believes something more urgent needs to be addressed: fighting for users’ copyright.

This is also what most frustrates creators today: after all the effort I put into prompting and refining a large model to write a song on a streaming platform, could contractual ties at the platform mean that the copyright does not even belong to me?

“Before discussing copyright, we first have to recognize users’ ‘right of authorship.’” Jiang Tao’s team makes its position explicit in the product: however you generate the work, copyright belongs to the user; the platform does not take it. In this era, users vote with their feet. Traffic will go to whichever platform protects creators’ hard work and helps them monetize it.

If Jiang Tao has chosen to break through by building an ecosystem, ACE Studio co-founder Zhao Wenxiao has made a classic move: changing the terms of competition.

History has already offered an answer to the question of how to resist closed-source giants.

“Think back: Midjourney was closed-source, but once Stable Diffusion became open-source, the entire ecosystem exploded, producing countless control plugins.” Zhao Wenxiao is following the same approach. ACE has released all the weights of its ACE 1.5 model as open source, letting users train their own Lora models and connect ComfyUI and their workflows.

If large models lock in a particular understanding of music, then dismantle the underlying arsenal and hand its parts to everyone. Overwhelming the isolated island of closed source with a vast ocean of open source: this is the most impressive card Chinese AI players have played under the shadow of the giants.

Low-End Dubbing Runs on Volume; Extreme Emotion Still Depends on People

Beyond music, AI speech synthesis, or TTS, is another heavily disrupted field. Many people think: with translation-and-dubbing software this powerful, and even models that generate audio and video together, are traditional dubbing companies serving international markets about to go out of business?

Song Kaifa, COO of Sudong Technology, which uses AI dubbing to serve overseas markets, punctured this technology-is-everything illusion.

First, a large model’s simultaneous audio-and-video generation is certainly exciting, but it applies only to new content generated from scratch. What is waiting to make money worldwide today is an enormous stock of existing Chinese short-form dramas, animation, and games expanding abroad. To earn dollars overseas, they need to turn northeastern Chinese accents into idiomatic southern Indian languages or German. In this area, repeatedly drawing random outputs from large models, as if playing a social-commerce lottery, simply does not work. It still takes specialized AI localization to tackle the problem.

But Song Kaifa described a crucial phenomenon: in this AI tsunami, lower-end voice actors really have lost their livelihoods. Yet highly experienced voice actors at the top of the pyramid have become more sought after.

Why?

Because AI cannot handle a person’s instantaneous loss of emotional equilibrium.

“Those extreme emotions that come in a very short space of time—‘I want to cry, I want to make a scene, I want to argue’—are simply beyond AI within 0.1 seconds or 1 second. Only people can deliver that brief emotional eruption.”

On that basis, Song Kaifa’s team has established an uncompromising rule: “If you have not said it, I will not dub it.” If someone brings in a celebrity’s voice and asks for it to be cloned to say something that celebrity never said, the team rejects the job outright. This is not just a boundary; it also defends the distinctive rights of high-level voice professionals.

Machines handle the inexpensive buildup; people deliver the expensive climax. That is the reality of AI dubbing.

If Everyone Will Be Replaced, Why Make Children Practice an Instrument?

As tools and business models race ahead, the pressure inevitably reaches education.

When a single sentence can generate a melody, should parents still push children to take graded music exams and learn one instrument—or several?

This anxiety, which cuts to the heart of the matter, was put to department director Duan Ruilei of Zhejiang Conservatory of Music.

“Good things are always made by smart people doing the slow, painstaking work.” Director Duan’s reply was unequivocal.

Having children learn an instrument in the AI era is not about competing with machines on how fast they can play without a wrong note. That would invite humiliation. The real purpose is to develop taste and aesthetic judgment.

This was also one of the clearest insights of the session. If AI can produce 100 tracks for you in one second in the future, what will enable your child to pick the most beautiful, most soulful one immediately from a vast ocean of computational junk?

Ears that have never endured tedious fingering practice or wrestled with harmony cannot pick out truly good music. This is, in fact, the paradox of AI as a whole: only people who genuinely understand a field can use AI to create something that surpasses AI.

It is like an online lesson: however clearly it explains things, it cannot replace a living person.

When you play a scale incorrectly, your teacher may offer a reprimand or a reassuring touch. “Education is one tree shaking another. I do not think today’s AI can replace that embodied companionship and emotional presence.”

All the panic, in fact, comes from treating AI as a “person.”

But this conversation tore away that filter. AI is always only a tool. It raises everyone’s baseline, letting even people who cannot carry a tune experience the pleasure of creating. Those who can break through the ceiling, however, are still ordinary people who have developed good taste through painstaking work—and who can break down crying within 0.1 seconds.

More Details from the Conversation

Unique Awards · Hangzhou AI WEEK Trends Roundtable Panel

“Music and Audio: AI Technology and the Industrialization of Content Production”

Guests:

Jiang Tao — CTO, Ziyou Liangji

Zhao Wenxiao — Co-founder, ACE Studio

Song Kaifa — Co-founder & COO, VMEG

Duan Ruilei — Director, Department of Music Engineering, Zhejiang Conservatory of Music

Moderator: Lu Xiaoxu — CEO, Xiaoxu Music

Lu Xiaoxu: I am Lu Xiaoxu from Xiaoxu Music. We run a specialist AI-music media account and have grown alongside the tremendous developments in AI technology over the past two years. I see many companies making comic-style dramas here today, and I wonder whether you have explored AI music, which is a very important part of that process. We have four guests onstage. First, could each of you briefly introduce yourselves and explain which part of AI audio your product or team works in? Let us start with Jiang.

Jiang Tao: Hello, everyone. I am Jiang Tao from Ziyou Liangji. Our product is Yinchao, a Chinese-developed music foundation model. Just before the Spring Festival, we released Yinchao 3.0. Across musicality, audio quality, arrangement, melody, and other dimensions, it is at a T0 level nationally. You can try it by searching app stores for Yinchao: the Chinese character yin means “sound,” and chao means “tide.” Our whole team shares a belief: we hope everyone can use music to record their lives and emotions, and we hope our product can help people fulfill that long-held wish.

Zhao Wenxiao: Hello, everyone. I am Zhao Wenxiao from ACE Studio. We are an AI-native music-creation tool with more than 1 million users now. They include Grammy winners as well as ordinary music enthusiasts. We recently open-sourced a music model called ACE 1.5, with all its weights available for download, so you can take a look. We are also developing a new project called ACE Music, a more open platform where people can create with a variety of music models and share their work.

Song Kaifa: Hello to all my industry colleagues. I am Song Kaifa from Chengdu Sudong Technology. We are a company serving international markets, primarily earning US dollars. Many people here today are probably our collaborators or partners. Our main work is using AI to make voices sound more human. Dubbing into major foreign languages for Chinese short-form dramas going abroad, film and television production, and marketing is currently our principal business. We are also here to learn. Overseas users account for 99% of our users, including people in India, Germany, Italy, and Thailand. Chinese short-drama companies expanding overseas are also a major user group.

Duan Ruilei: Hello, everyone. My name is Duan Ruilei, and I come from the Department of Music Engineering at Zhejiang Conservatory of Music. Like Mr. Lu, I work in digital music creation. The name “Department of Music Engineering” may sound unusual at a conservatory. It is a department dedicated to teaching and talent development in digital music.

Discussion One: Changes in the AI-Music Copyright Market in 2026 and Creators’ Rights

Lu Xiaoxu: Our four guests cover the different sides of AI audio: music foundation models, all-in-one AI-music workstations, voice services for overseas markets in Mr. Song’s case, and music education in Director Duan’s case. They represent different directions, so our questions will vary by field. First, I would like to ask Jiang and Zhao: what has been the hottest subject in AI music this year, 2026? AI-music copyright. There are more and more creators of AI music, including people in the audience who may not have formal musical training but can still write songs with AI. How do they make money from the resulting music? With millions of AI music works appearing on Suno every day, and copyright arrangements across platforms in China gradually developing, how do you, as representatives of AI-audio tools and foundation models, view changes in this market in 2026?

Jiang Tao: Before discussing copyright, we should take an earlier step and ask whether we recognize users’ right of authorship over music. I think we agree on that. Users add their own Prompt, ideas, and imagination, and keep refining the result. They have invested effort and care. So I think their authorship of that AI music work should first be recognized. Then we can discuss copyright. Copyright may involve some historical mountains we can never scale, but after not going around those mountains, we can still see that consumption of AI music is extremely high. This is especially true of derivative fan works. For example, I saw a derivative fan concert based on A Record of a Mortal’s Journey to Immortality: ten songs for one character, making a concert nearly 60 minutes long. Some of the works were quite good. They have high play counts and many comments on QQ Music and NetEase Cloud Music, showing that users like this kind of work. If a work has authorship and users like it, how can users receive reasonable income? That is what platforms need to figure out. If streaming platforms can satisfy both income and creative needs, creators will certainly gravitate toward them. We see some streaming platforms refusing to pay out for AI works while others do pay. Platforms that pay have recently seen 80% growth in monthly active users. That is plain to see: users vote with their feet.

Another point is that platforms such as Yinchao and ACE recognize users’ authorship and the associated range of copyrights. Some platforms have been tied into arrangements with major rights holders, causing creators’ copyright to be separated out or reassigned. This is the dilemma all creators face: “Should I create on this platform at all? If I make something good, its copyright may not even belong to me.” Creators will naturally move toward platforms with more accommodating copyright arrangements and stronger protections for creators. Yinchao’s license, for example, explicitly says that all copyright in creations belongs to users and that the platform does not take ownership. We can also issue logs and supporting documents showing when a user entered a particular Prompt to create a work and that authorship belongs to that user. That provides a form of rights protection.

Zhao Wenxiao: Jiang has already covered this very thoroughly. Copyright initially existed primarily to protect creators’ ideas and the results of their work, because creation takes a great deal of time and effort, whereas copying is easy. Earlier, copyright income might have meant receiving a share of CD and cassette sales; later, streaming meant being paid according to song plays. Both are essentially about protecting creators’ rights. ACE was not originally a fully end-to-end AI-music generation tool. It was a creation platform that helped creators generate content using tools. I certainly think things into which creators have put their hard work need protection.

There has also been a lot of discussion recently about using data to train models. We started relatively early: we previously had AI singing voices. Our approach is partly to purchase data ourselves or hire people to record it, and partly to work with well-known instrumentalists or singers to release virtual sound libraries or virtual instruments bearing their identity. When users use those resources on the platform, we ultimately settle payments according to usage and give the artists a percentage. So in the future, if a type of model is distilled from an artist’s style or uses a Lora based on that style, a share may also need to go to the artist during use. Jiang just mentioned that mainstream streaming platforms may not currently accept AI-music copyright, but I think this will open up increasingly. ACE Music, which I mentioned earlier, is also a relatively open platform. I think anything embodying people’s creative effort can earn a return.

Lu Xiaoxu: On copyright, users really vote with their feet. As long as it sounds good, they do not mind whether it was created by a person or AI. A lot of data already suggests that songs created by AI may receive more likes from listeners.

Discussion Two: Do We Still Need to Learn Music and Instruments in the AI Era?

Lu Xiaoxu: My next question is for Director Duan. As a creator on WeChat Channels, I get the most comments from students and teachers. Many say that AI makes producing songs so easy now: do we still need to study music? Do we still need to learn instruments? How can we use AI to support learning? I imagine this is also a huge challenge in your teaching. How do you understand this question?

Duan Ruilei: This has been a serious question facing us over the past two years. Here is how we think about it: should we first understand music as a noun or as a verb? I lean toward the verb. If it is a verb, it means participating in music and engaging with it, whether through appreciation or creation. I believe AI makes both more accessible. On that basis, whereas we used to discuss how to stop students from cheating with AI, we no longer think the so-called “cheating” issue is really an issue. What matters is using it well. So to ask whether people still need to learn music, let us return to the act of engaging with music.

Our department still primarily trains people in music creation and production. We keep emphasizing one idea to students: “Good things are always made by smart people doing the slow, painstaking work.” However advanced AI becomes, the quality of music generated by different people still varies widely. Music also differs from visual work, where you can look at a hundred pieces at a glance. It is an art of time: if you generate 100 minutes, you must spend 100 minutes of your life listening to it. So musicians will always have the task of putting in the painstaking work, combining AI with their musical understanding and taste to recommend or produce high-quality music for others.

Lu Xiaoxu: Director Duan just mentioned something very important: aesthetic judgment. My own child is eight and in the third year of primary school. Should they study music? Absolutely, and the earlier the better. Learn as many instruments as possible. Why? Not so that we can create music to compete—PK—with AI, but so that we develop our own taste. If AI can give you 100 tracks in a minute, can you quickly choose the one you think is best? Aesthetic judgment needs to be developed from childhood. It is very important.

Discussion Three: Balancing Overseas Dubbing with CV—Voice Actors’—Rights

Lu Xiaoxu: My next question is for Mr. Song. What has been a hot news story lately? Many well-known voice actors have issued statements defending their voice rights and opposing illegal AI voice cloning. I have seen CCTV report on it too. As an expert in AI voice services for overseas markets, how do you balance voice actors’—CVs’—legitimate rights with the future development of AI speech? What approach do you take?

Song Kaifa: We have more than 2 million users worldwide and mainly focus overseas. I will first discuss overseas opportunities related to music and then come back to China. Many countries do not have digital-content supplies as developed as China’s or North America’s, including many countries in the Middle East and Europe. For instance, some Black rappers are very good at rapping but want to change their accent to standard American or Australian English so their music can travel further. They pay us for translation and dubbing, sometimes quite substantial amounts. This includes translating music in Hebrew into English while retaining its original melody. Someone might pay US$100 or even US$1,000 for a minute. So we see many opportunities overseas in this area.

Coming back to China, we also see voice actors experiencing ambivalence and conflict over AI. It is comparable to translation: AI has replaced a lot of basic translation work, but high-end, experienced translators remain expensive and valuable. With dubbing, China now uses AI extensively to take short-form dramas overseas, primarily to reduce costs and improve efficiency. People often ask whether I can produce 1,000 or 5,000 minutes of translated dubbing a day, turning Chinese into Spanish, Thai, or Arabic for overseas audiences. That is somewhat lower-quality, volume-driven work. But major overseas clients we encounter, such as large Malaysian digital-media companies wanting to dub animation into Indonesian, listen to AI and find it inadequate. They are willing to pay a high price for human dubbing.

So entry-level voice actors without a strong personal style will indeed be hit hard in this industry. But for experienced professionals, our business rule is: we translate and dub only things you have actually said; we will not dub things you have not said. If a user wants a celebrity’s voice cloned to say a line, and that celebrity never said it, we will not take the job. This protects high-level voice professionals’ rights. I have also spoken with many directors. Even after using international models such as ElevenLabs or Chinese models from MiniMax, iFLYTEK, or Tongyi Qianwen, high-quality content still needs professional voice actors. Those extreme emotions within a tiny space of time—“I want to cry, I want to make a scene, I want to argue”—are beyond AI within 0.1 seconds or one second. Humans can manage those brief outbursts. In the future, experienced voice professionals can work more with organizations to understand and protect their rights, which can also help the whole industry develop. Most AI is still used for high-volume short-drama work. The directors at the heart of major feature-film production will usually still need human voice professionals rather than AI.

Lu Xiaoxu: That is very well put, and I feel it too. We have long produced music and dubbing for games. At present, 80% of game dubbing still needs real people, because AI dubbing remains somewhat lacking in strong emotional expression and characterization. I think it may grow stronger as technology develops, but the emotion people bring is irreplaceable.

Discussion Four: The Advantages of Chinese AI-Music Tools and the “Competitor” Ecosystem

Lu Xiaoxu: My next question is again for Jiang. Whenever we discuss AI-audio tools, we inevitably think of Suno. As a leading Chinese AI-music model, how do you benchmark against and compete with Suno—or scale that mountain? Suno has also launched a Chat agent feature. Where do you see the advantages of our domestic AI-music products?

Jiang Tao: Suno is indeed an industry leader, both in model performance and product features. The Audio feature you just mentioned was not actually introduced by Suno first. Riffusion, which was acquired by Google, pioneered that approach, and our small overseas-market product hitto.ai also uses it. But I want to stress that we are not in a competitive relationship with Suno. Suno’s user-facing product increasingly targets professional musicians, and the features in Studio are becoming more complex. But music serves the general public as well as professionals, and there are other use cases. This is an enormous market, with many audiences’ music-production and consumption needs still unmet. Film, television, and games, for example, still offer vast room. So we see no need to fight Suno head-on in creation tools. There are many other user groups we can serve. Working together to make the market bigger and stronger is what matters most. In today’s technological iteration, although we hope to reach T0 performance, we focus more on maximizing returns from the Yinchao model in IP music, film, and television. Everyone should do their best in their strongest areas. As a CTO, I still hope underlying models keep improving so we can learn technical ideas from them. From Suno’s Bark model opening up an era in ’23 to the present, Suno has kept leading. But whether leadership maximizes returns is another question, because every country has geographic and cultural differences. Meeting the emotional needs of cultural audiences we understand produces different returns. So we hope Suno keeps getting better, and we also hope to keep progressing on our own path.

Lu Xiaoxu: A question for Zhao: ACE Studio is primarily an integrated music workstation for professionals and semi-professionals, so comparison with Suno still comes up. Although ACE’s international influence is growing and it increasingly resembles a global DAW music-production product, what is ACE Studio’s strategy on these two points—benchmarking against competitors and internationalization?

Zhao Wenxiao: We noticed this very early. Suno’s initial model was called Bark. When we built ACE, we were pursuing a high degree of control, with AI voices and many small tools to assist creators. At that time, Suno simply took a Prompt and produced a song end to end. As time has passed, Suno has inevitably moved toward greater controllability, while we have also been developing a music foundation model. So ultimately there really is competition between us. That is unavoidable.

Our strategy is this: first, no single company will dominate the entire model landscape. Look at today’s large language models. OpenAI originally dominated, but now many powerful models have emerged, such as OpenClaw. Different models have different characteristics. In the future, some might be better at classical music, others at rock or pop. Creators can combine models according to their needs to produce the best work. Another point is our open approach. Think back to closed-source Midjourney and the open-source image model Stable Diffusion. The flourishing ecosystem brought about by Stable Diffusion’s release was something Midjourney could not match. Tens of thousands of Lora, ControlNet, and other control plugins subsequently emerged. Combined with ComfyUI workflows, they can produce results beyond a single closed-source model. ACE now follows this strategy too: we have open-sourced all the weights, support people training their own Lora models, and received immediate support from the ComfyUI community. Finally, a real creative work still depends on people. People have their own cultural understanding, whereas a model’s understanding is fixed. Music ultimately advances through everyone creating in many different ways.

Discussion Five: The Impact of Visual Models’ Simultaneous Audio-and-Video Generation on Audio

Lu Xiaoxu: My next question for Mr. Song may be a little pointed, because there are many comic-style-drama companies here. Video models such as Sora and Kling, which generate sound and pictures together, affect us substantially. Our company used to provide sound effects and music. If the model produces both, there is not much left for us to do. Might that also affect tools such as TTS—text-to-speech? How can you avoid the possibility that future visual foundation models absorb audio as well?

Song Kaifa: We see a great deal of content worldwide that has not circulated because of language barriers. If voice professionals feel constrained in Chinese-language work, they can move toward multilingual work. There is huge demand for dubbing in Cantonese, Tibetan, and “Xinjiang language,” for example, but good voice actors often do not know those languages. Returning to your question, we encounter a vast amount of already-produced content worldwide that needs translation and dubbing. YouTube, for example, supports multiple languages, and many major creators use our TTS services to dub their content for global audiences, with more lifelike pronunciation.

Models such as Sora and Kling that generate audio and video together target incremental content—new content—and can produce dubbing directly. But there is already an enormous stock of existing content. Serving it requires a good understanding of local culture, translating local jokes, and refining speech with appropriate interjections and tone. Large models still struggle to meet those needs. So in the short to medium term, their impact on translation-and-dubbing tools like ours is very small. Also, for content whose sound has been conceived in advance and needs precise control, large models may not meet the requirements. Repeatedly sampling outputs becomes expensive. Hiring a person or using TTS may be better; producing the components separately can actually cost less. Large models greatly improve content production, but their impact on traditional needs is not that large. There are still substantial opportunities.

Lu Xiaoxu: That is very clear: one concerns existing content and the other new content. Kling’s simultaneous audio-and-video generation mainly addresses new content. But for the vast amount of video already made, using traditional human dubbing again would cost too much, so there is more reason to use Mr. Song’s product.

Discussion Six: Music Education and Future Trends in the AI Era

Lu Xiaoxu: My next question is for Director Duan. We discussed how students should learn, so now we must ask how teachers should teach. In April, the education bureau in Shanghai’s Yangpu District asked me to provide AI training for more than a hundred primary- and secondary-school music teachers from across the city. I feel quite a lot of pressure. How should teachers use AI to support and improve music education? What plans or views do you have?

Duan Ruilei: Let me offer my understanding. Music learning falls into two categories: acquiring knowledge and training skills. Because of large language models’ inherent strengths, much knowledge-based learning has already been restructured. A traditional classroom semester’s knowledge content may be covered very quickly. But for skills training, I have not yet found a particularly good product that can replace a teacher supervising practice. Bowing angles, the force applied to keys, and training the fingers’ various functions, for example, still lack a particularly good tool. To extend that thought and return to education itself: education should fundamentally be one tree shaking another, one mind touching another. During learning, a teacher’s words and example, emotional presence, companionship, encouragement, and criticism are things I have not yet seen AI able to replace. Human beings remain human.

Lu Xiaoxu: When a child learning an instrument plays a wrong note, a teacher’s touch communicates human feeling or encouragement. I think that remains very difficult for current embodied intelligence to replace. But there are opportunities at the level of tools and knowledge. There are not yet especially good teaching-support products, which is an opportunity for entrepreneurs. How to use AI to learn music well is very important for teachers and students alike.

Lu Xiaoxu: Finally, could each guest give a brief conclusion? Most people in the audience today are not music professionals or are AI enthusiasts. How should everyone explore AI music in 2026? And what are the overall development trends for AI audio after ’26?

Jiang Tao: People developing music foundation models share a view: in three to five years, most new music people consume may be made by humans with AI assistance. There may not be much music made 100% by the old methods. But the most creative and imaginative portion will certainly still come from outstanding human creators. AI can already handle the middle tier. What people will increasingly pay for is whether music meets their emotional needs. I have a deep belief in AI. I have watched it go from singing chaotically to an emergence of intelligence that produces beautiful singing. So I believe that in 2026, what everyone needs to learn is how to advance and grow alongside AI. It is almost an all-round performer; much depends on the use case and emotional expression a creator needs. Everyone can be 100% confident that AI can do it. The time and compute costs may currently be somewhat high, but they will gradually fall. This year is a very important learning process.

Zhao Wenxiao: I used to work in technology, and you can see the trend in AI coding. Writing code used to be complex: configuring environments and learning theory took a long time. Now you can state your requirements to an Agent, and it can build the functionality. Music shows a similar trend. Producing it once required downloading a complex DAW, sound libraries, and effects, and learning extensive engineering knowledge. Future creation may be based primarily on taste: have AI generate something, then refine and correct it step by step according to your ideas. The boundary between creation and consumption will become increasingly blurred. While listening, you may think of an improvement and immediately move into creating. Creation will become very interesting, and more works in different styles will emerge.

Lu Xiaoxu: Your two products lean toward 2C, while Mr. Song works in 2B. Could you briefly discuss the development trends for voice and dubbing after 2026?

Song Kaifa: Translation and dubbing are internationally called localization—Localization. When translation and localization are efficient and inexpensive, they can help culture travel, taking Chinese stories, martial-arts tales, and cultivation fantasies to Belt and Road countries and beyond. Telling our own stories deepens mutual understanding. From last year to this year, overseas orders have kept increasing. Foreign organizations also recognize China’s strong content-production capabilities and come directly to us for localized dubbing in other languages. Demand for dubbing that engages deeply with national cultures will grow. Large multinationals training employees, for example, are gradually moving from uniform English delivery to personalized audio in local languages. AI has created more language-localization demand. There are many future B2B opportunities, and many translation-and-dubbing organizations now work overtime every day. This is a very good opportunity.

Lu Xiaoxu: A final question for Director Duan: how should parents here help children preparing to apply to music programs in the AI era? What plans are there for university–industry collaboration?

Duan Ruilei: As AI technology and the economy develop, some people will gradually stop treating music simply as a way to earn their daily bread and genuinely return to the art itself. Those who truly love music, are good at it, and are suited to it will gradually be the ones who remain in the industry. For university–industry collaboration, universities’ most important task is educating people. We attach great importance to working with companies at the technological frontier. Companies feel much greater urgency about R&D than universities do. If students are interested in their research directions and products, we actively encourage them to enter industry. I also invite our friends in business to come to Zhejiang Conservatory of Music to recruit. We have many talented students and look forward to more collaboration.

Lu Xiaoxu: Let me finish with a couple of thoughts. Music is an important form of artistic expression that is often overlooked. Everyone can see that competition in AI video has become extraordinarily intense, yet there are not many AI-audio products. I hope teams with product-development capabilities will put more effort into specialized areas of AI audio. As the saying goes, “A good show cannot come off without sound.” Our guests represent academia, voice services for overseas markets, music foundation models, and workstations. I hope you will continue talking afterward. Thank you to all the guests for sharing and to everyone for attending today’s forum. Thank you!

Originally published by Unique Research on Unique Research Substack on April 12, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research