跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

After Sora 2 and Nano Banana Pro, the Barriers to Creation Collapsed—and the Ceiling Soared

Original · Unique Research · 2025-11-25

Editorial note: This full historical report and panel transcript preserves the original speakers' claims, including product rankings, adoption figures, roles, service availability and forecasts; these are not independently verified present-day claims. The transcript says Sora 2 launched in November, but OpenAI's original announcement is dated September 30, 2025 (https://openai.com/index/sora-2/). The speaker's wording is retained with that discrepancy disclosed. The scale of Wang Chuan's three earlier platforms has no specified metric and must not be read as a count of enterprise customers. The second-half-of-2025 real-time-interaction statement remains a historical forecast, not a confirmed release.

If, for the past hundred years, admission to the world of moving images has been held by a select few, then in 2025 that ticket is being reprinted—at a scale of one for every person.

The topic of this panel at the 2025 Unique Bloom in Beijing was straightforward: Who will be the creators of the AI era? And where will large video models take the content industry?

The four guests onstage represented four of the most powerful currents in this wave:

Richard Wu Weijie, Senior Vice President of Zhipu AI; Sun Weizhe, head of AIsphere Technology's B2B business; Wang Chuan, Vice President of ShengShu Technology/Vidu; and Chris Su Guoli of Sand.ai. The moderator was Zhao Liang, a partner at Unique Capital.

They were not merely discussing trends. They were unpacking the question before the audience: Now that Sora 2 has raised video-model capabilities to a new level, how exactly will creation move from professional studios into everyone's daily life?

I. Four Aha Moments That All Point to the Same Thing

Moderator Zhao Liang asked a direct question: When did AI strike a chord with each of you and persuade you to enter the field?

Chris did not point to a flashy demo. He described a deeper insight inspired by GPT-4o:

Models are evolving from toolboxes into built-in capabilities.

Building a product used to be like assembling blocks: a language model, an image model, workflows, and plug-ins, all connected piece by piece. What stunned him about GPT-4o's image generation was that complex operations that once had to be connected through workflows had been folded into the model itself.

As a result, "the model is the product" is no longer a slogan. It has become a challenge for R&D teams:

Should this function be written into the product, or trained into the model?

Wang Chuan's defining moment came with Sora 1. His reason was simple:

Video finally stopped looking like a concept reel and started looking like reality.

He saw more than convincing realism. He saw disruption across industries: interactive entertainment, e-commerce, education, gaming, and almost every other sector could have new doors opened by video models.

For him, the original reason for joining this wave was to apply frontier technology to real production—not to put it on display, but to make it useful in customers' factories, classrooms, and livestream studios.

Sun Weizhe had two such moments. ChatGPT first drew him into AI. More importantly, after a transformation-effect template was launched, PixVerse gained users on the ten-million scale in a single month.

What he saw in that viral growth was this:

The real inflection point is not how powerful the technology is, but how low the barrier becomes.

ChatGPT was not the most powerful model, but it made AI usable by ordinary people. Packaging prompts into templates turned video AIGC from an expert's toy into a language everyone could use.

At AIsphere Technology, they called it video AIGC's GPT moment.

Richard of Zhipu AI did not have a romantic flash of inspiration. He made a practical judgment:

This was already a market capable of moving forward on its own.

In 2023, before Sora had been released, Zhipu AI had already achieved strong results in both technology and commercialization.

He jokingly summarized his decision-making method this way: If you are returning to technology, work on the most cutting-edge large models.

It sounds casual, but behind it was a rational bet made after weighing the trend, the team, and the ability to deliver.

II. When Models Begin to Fill in the Blanks, the Creative Relationship Is Rewritten

Chris offered a vivid judgment: Sora 2 gave video models a different way to solve the problem.

The previous path was to chase metrics—resolution, motion quality, camera stability, temporal consistency—as if taking an exam.

Sora 2 is more like a model learning to understand you. You only need to describe your intent; the model fills in the details and can even devise the presentation on your behalf.

The underlying change is that people move from executors to directors.

A director does not need to carry the camera, set the lights, edit, or score the film. The director only needs to keep giving feedback:

Is it good or not?

Make this part a little faster.

Do not make it so sentimental.

AI turns the director's intent into cinematic language.

In other words, the better the model becomes at filling in the blanks, the more the human role shifts toward high-level decisions.

Creation moves from physical execution to aesthetics and choice.

And everyone has at least some aesthetic judgment. What people lacked before was a camera capable of understanding them.

III. Who Exactly Is an AI Creator? The Answer Is Broader Than You Think

This part of the discussion was especially interesting because it expanded creators from a small circle into a new identity spanning society.

Richard divided Zhipu AI's creators into three ecosystems:

Business creators in large organizations: ordinary employees working in areas such as urban transit, energy, and emergency response. They use models every day to improve services and are, in essence, creating new value. We once called them staff members; now they are people redefining their jobs with AI.

Engineering creators in internet and technology companies: leading Chinese internet companies are engaging in collective creation within the model ecosystem. Teams are using models to rewrite their businesses, improving user experience and overall efficiency.

Independent developers: Paid subscriptions to the GLM Coding plan surged in the months after launch, showing that code is becoming a new language of creation and that its barriers are falling.

More strikingly, he mentioned a one-person-company entrepreneurship program.

The aim is not to turn everyone into a startup celebrity, but to let more people try using AI to create a new livelihood.

Even if only 1,000 of 1 million one-person companies succeed, and only 10 ultimately have investment value, the experience would still teach all 1 million participants how to work with AI, improve their employability, and reduce anxiety.

Sun's classification was more like a social history of creative life:

Super-creators are those at the leading edge. Like DJs in the multimodal era, they combine LLMs, TTS, and video models into new workflows. PixVerse gives them capabilities as well as a stage through film festivals and competitions. His story about a super-individual artist carried particular weight: after suffering a setback, one person relied on an understanding and love of art—and on PixVerse—to return to the professional stage. Technology did not encourage that person to withdraw from life; it helped them stand up again.

Ordinary creators are everyone. Upload an image and write a few sentences, and you can receive an end-to-end story complete with sound, subtitles, and a cover image. "Everyone is the director of their own life" used to be a slogan. Now it is becoming a real product capability.

Wang Chuan described AI creators with three keywords: efficiency, combination, and globalization.

Finish a month's editing workload in one week;

combine multiple multimodal tools;

and think about global markets from day one.

Together, these three elements create a new production function for creativity—compressing time, tools, and market boundaries at once.

The boundaries of the AI creator are therefore becoming more like oxygen:

It is not an industry, but a way of life.

IV. Why Are AI-Animated Microdramas Booming? Because They Occupy a Higher Beta Segment

The discussion of AI-animated microdramas also took a broad, flowing perspective.

Chris said that AI-animated microdramas are an even higher-Beta subcategory within the already high-Beta microdrama market.

The reason is straightforward: microdramas have limited budgets and cannot stage vast, surreal spectacles;

AI-generated images and animated formats, by contrast, are naturally suited to letting imagination overtake production constraints.

That creates a new possibility: an era of super-creators in which one person can make an entire film.

With video models undergoing paradigm-level changes every 3–6 months, and AIsphere having iterated through 8 model versions in two years, future super-creators may use a model stack to produce realistic works that rival traditional studio production.

The next explosion in microdramas would then be not growth in quantity, but a breakthrough in the threshold of production quality.

Wang Chuan was more pragmatic: Vidu is already used by 70% of leading studios in the AI-animated microdrama sector.

He identified its key capabilities:

Duration and pacing control, adjustable from 1 to 10 seconds

Strong character-expression performance, with more natural crying and smiling

Stable consistency, so characters' faces do not change

More importantly, he described two schools of workflow:

Image-to-video: script → storyboard → image generation → animation

Reference-to-Video: first build libraries of characters, props, and settings, then generate video from a single sentence

The second approach is more like digital live-action production. It is upgrading AI-animated microdramas from assembling pieces to shooting a film.

He ended with a point of suspense:

The decisive variable in AI-animated microdramas may not be better visuals, but a new entertainment form—real-time interaction.

One-way dramas could become living content in which viewers participate, make changes, and converse with characters.

V. When Will Real-Time Interaction Arrive? The Answer Lies in the Architecture

This exchange was the most futuristic part of the entire session.

Chris explained why Sand.ai chose an autoregressive path: Diffusion video generation normally requires waiting for every frame before playback, while autoregressive generation can create and play the video continuously, like flowing water.

When the granularity is reduced from 1 second to 1 frame, real-time interaction moves from concept to engineering reality.

He even offered a very specific time marker:

Early signs will appear in the second half of 2025.

This is not simply a prediction that something will happen someday. It is an advance bet, through an R&D roadmap, on the shape of a future market:

Architecture determines the form of interaction, and the form of interaction determines new content industries.

Put differently, the model architecture chosen today may become the physical laws of tomorrow's entertainment world.

Conclusion: Creation Is Moving from Effort to Second Nature

Beneath this panel, a powerful consensus seemed to flow:

Video models are not simply producing video more capably. They are becoming a medium that better understands human expression.

As models acquire more native capabilities, people can redirect their energy from how to make something toward what to make and why;

as barriers continue to fall, creation moves from a profession for the few to an everyday activity for the many;

and when real-time interaction arrives, content will no longer be a finished product, but a living entity that can grow together with its audience.

Perhaps, a few years from now, we will look back at today's AI-animated microdramas, transformation effects, and short videos as we now look back at the earliest short-video platforms:

They were all signals of one thing—

creative agency is being distributed more widely.

It is not AI replacing creators.

It is AI returning the power to create to more people.

More Details from the Conversation

01. Opening and Guest Introductions

Zhao Liang: Hello, everyone. I am Zhao Liang from Unique Capital and the moderator of today's first panel. I am delighted to welcome Richard Wu from Zhipu AI, Sun Weizhe from AIsphere, Wang Chuan from ShengShu Technology, and Chris Su from Sand.ai.

Some people have said that the moderator may be the least necessary person on today's panel. Someone suggested that the guests question one another, which would certainly be lively. Every guest here works at a company building large models, general-purpose large models, multimodal large models, or video models, and their companies may have competed on many fronts. That is all right; today's panel will be relatively peaceful—"Peace," if you like. But consider this fair warning: perhaps we will invite everyone back for another event and let the guests question one another. That might be even more interesting.

Let us begin formally. First, please introduce yourselves briefly. Since Sora 2 was released in November, all of you have made new moves, and different companies have released new models. Let us start with Richard.

Wu Weijie: Hello, everyone. I am Wu Weijie from Zhipu AI, where I am mainly responsible for commercial implementation. Our company grew out of Tsinghua University and currently serves more than two million enterprises and developers worldwide. We are a startup committed to independently developing foundation models and maintaining autonomy and control across the full technology stack. I hope we will have opportunities to communicate and collaborate with the guests and everyone here. Thank you.

Sun Weizhe: Hello, everyone. I am Sun Weizhe from AIsphere. AIsphere is a company that builds large video models. Our product PixVerse, called Paiwo AI in China, has more than 100 million consumer users worldwide, and our B2B business is also growing rapidly. I lead the B2B business. I have been an entrepreneur in this field for quite some time, and Richard and I have known each other since our days at Feishu.

Wang Chuan: Hello, everyone. My name is Wang Chuan, and I have worked in AI for many years. I previously worked at several major internet companies, primarily on AI at Tencent, NetEase, and Baidu. I successively built about 3 AI open platforms serving enterprise customers, described in the source as operating at the tens-of-millions or over-one-hundred-million scale, without a specified metric. I now lead enterprise services at ShengShu, bringing our model capabilities to external customers. You may know our product Vidu. We are also a multimodal company focused on video, and we develop our own models. Like Zhipu AI, we also have roots at Tsinghua University. We serve tens of thousands of developers and more than one thousand enterprise partners worldwide. Thank you.

Su Guoli: Hello, everyone. I am Su Guoli from Sand.ai. Sand.ai develops models focused on AI video. Our long-term vision is to democratize AI for humanity and build a more consumer-facing (To C) business. Until now, we have placed greater focus on model R&D, but after the release of Sora 2, we may gradually accelerate the implementation of our consumer business.

I am responsible for product work at the company. I have been an entrepreneur since 2018, and I am a serial founder. Before joining Sand.ai, I also created an AI product at Vast: Tripo, an AI 3D tool for creators.

02. The Aha Moment That Led Each Guest into AI

Zhao Liang: I would like to ask a personal question. When you first joined the wave of AI creation, was there a particular Aha Moment that amazed you and made you commit to the field? Looking back, is that original motivation reflected in your product today? Chris, please go first.

Su Guoli:

My earliest sense of this AI wave certainly came when ChatGPT was first released. But the milestone most closely connected to my own field arrived early this year, when OpenAI released GPT-4o image.

GPT-4o's image-generation ability made a profound impression on me because it felt very much like a Vertical Agent. It can internalize within a single model many complex image-editing capabilities that previously required workflows, multiple tools, and calls to language models.

When I saw it, I realized that this was a milestone in a major trend: the complex intelligence centered on language models has the potential to be integrated into multimodal models. That had a major impact on us. While training video models, Sand.ai spent a long time studying how to unify an Autoregressive architecture with traditional Diffusion generative models. GPT-4o is probably moving in a similar direction, integrating a language model and image-generation model in training from Day One.

This profoundly changed the definition of "the model is the product." Previously, even a product such as Midjourney needed only a simple Discord interface. But GPT-4o made us realize that for every product feature we want to design, we should first ask whether data processing and training could make the capability native to the model.

From that point on, we ran many highly imaginative experiments. Some time ago, we released the Gaga-1 model, the world's second video model capable of generating synchronized audio and visuals. Others in the same cohort included Sora 2 and Alibaba's Wan 2.5. All of this benefited from that series of experiments.

Wang Chuan: We also greatly respect the approach Chris just described. At ShengShu, we primarily apply a DiT (Diffusion Transformer) architecture.

For me personally, the most exciting moment still came when Sora 1 was released after the Spring Festival. Every major group chat was discussing it. I watched many videos and felt the released material had become almost indistinguishable from reality. Later examples, including the balloon-person advertisement and collaborations with Hollywood, convinced me that AI was steadily moving from language into an increasingly mature era of video and multimodal implementation.

I quickly spoke with customers in several industries and found that video or multimodal models could serve an extraordinarily broad range of sectors, including interactive entertainment, the internet, e-commerce, education, and gaming. Expectations for the model were extremely high. I therefore decided to join this wave. My original aspiration has always been to combine advanced technology with customers' production practices.

Sun Weizhe: I personally had two Aha Moments.

The first was also ChatGPT. I was stunned and felt I had to work in AI. At the time, I explored many applications, including working with leading MCNs and film and television companies on AI topic selection, scripts, and RAG, among other applications. But I found that if I operated only in that niche on the B2B side, I did not have much of an advantage.

The second moment came last October. AIsphere had launched a template feature that packaged prompts for users with one click, including a special transformation effect. PixVerse was not very well known before that, but this single effect attracted users on the ten-million scale in one month, largely through organic sharing.

This gave me an insight: the ChatGPT moment was not purely a technical competition. It was the democratization of technology that made participation possible for everyone. AIsphere's transformation-effect template likewise brought video AIGC into ordinary households for the first time, rather than leaving it as a toy for advanced users. This could be considered video AIGC's GPT moment. From then on, I felt AIsphere had tremendous room to grow in video AIGC, so I later joined the company.

Wu Weijie: It is also difficult for me to name a particular moment. After leaving ByteDance, I spent two years investing at a traditional company. When ChatGPT became popular, I considered returning to technology. I called a former subordinate who was still at ByteDance. He said, "Why keep working on the internet? If you are going to do something, work on large models." He suggested I research China's six leading large-model startups.

I also considered two companies in Shanghai, but our directions did not align after we spoke. It was late 2023, before Sora had appeared. Zhipu AI was one of the few companies in the market already pursuing commercialization and had achieved highly impressive commercial-order results in its first year. I found that remarkable. Taking everything into account, I thought it was a strong field and joined.

03. Defining AI Creators and Supporting the Ecosystem

Zhao Liang: Today's creators are no longer content producers in the traditional sense. They are people who use AI tools deeply—including for short videos, AI-animated microdramas, visual effects, and programming—to create new value. From your companies' perspectives, how do you define these AI creators? What specific support policies do you have for current markets, such as the booming AI-animated microdrama sector?

Wu Weijie: We divide our customers into three categories, corresponding to three ecosystems:

Large state-owned enterprises and urban ecosystems: We work with provinces and cities to form local ecosystems. One recent example is Hangzhou City Investment, where we empowered public transit, energy, emergency response, and other services. These ordinary business employees use large models every day to serve thousands of households, and we consider them creators as well.

Internet and high-tech companies: Most of China's Top 10 internet companies, including Meituan, work with Zhipu AI. This year, after Claude (Anthropic) stopped providing services in China, Zhipu AI benefited, and many users who had previously used Claude switched to Zhipu AI.

Independent developers: We introduced the GLM Coding plan in September, and it has grown very rapidly over the past two months.

Regarding the one-person enterprises and creators the moderator mentioned, we have indeed been studying the idea and are considering launching a one-person-company entrepreneurship program. Zhipu AI and its partners could provide a platform, venues, instructors, and a complete set of services. The goal need not sound grand. For example, if we serve 1 million entrepreneurs and 1,000 of them succeed, with roughly 10 ultimately receiving investment, that would be very good. Even if the rest do not succeed, the experience can still give them a better understanding of AI, improve their employment skills, and ease anxiety.

Su Guoli: Over the long term, we hope everyone will have the opportunity to become a creator. Every person has unique life experiences and something they want to express.

Sora 2 gives video models a different solution. Rather than focusing solely on the traditional metrics of video—even though it is already at the SOTA level—the model imagines the missing details and takes responsibility for the presentation. Users need only describe their intent simply.

This changes the relationship between people and creation: people become increasingly like directors, while AI handles the cinematography and presentation details. People only need to keep offering feedback: whether it is good, and what should be changed.

As for AI-animated microdramas, I see them as a higher-beta subcategory within the already high-beta microdrama market. Because budgets constrain microdramas, they struggle to produce large-scale spectacles such as The Wandering Earth. By using AI-generated imagery or animation, AI-animated microdramas can depict surreal scenes more easily.

But as video models improve—with a paradigm-level shift roughly every 3–6 months—a super-creator may eventually use models alone to produce realistic works that rival Hollywood blockbusters. At that point, the microdrama industry may see a much larger boom, and its production values will no longer feel like a quick, cheap substitute.

Wang Chuan: I define AI creators in two broad categories: individual entrepreneurs and institutions or studios. They have three characteristics:

An extreme pursuit of efficiency: content that once took 1–2 months can now be completed in a week;

they are mixers of multiple tools, using not only video models but also language, audio, and image models—hybrid users of multimodal tools;

and they have a global perspective, thinking about international expansion from day one.

In AI-animated microdramas, 70% of leading studios are now using Vidu for related content. Our Vidu model is particularly well suited to microdramas because it offers:

Second-level control: durations from 1 to 10 seconds can be controlled precisely;

strong performance in dramatic scenes, with effective control over facial expressions such as crying and smiling;

and high consistency.

AI-animated microdramas now use two schools of workflow:

One is image-to-video: script → storyboard → image generation → animation;

the other is Reference-to-Video, which we are promoting. It is more like live-action shooting. You first build libraries of protagonists, props, and settings, then upload a protagonist, prop, and background to generate video from a single sentence. A leading studio called Jiangyou uses this method extensively.

Finally, I believe the only major variable for AI-animated microdramas is a new entertainment form: real-time interaction. Today's dramas are all one-way outputs. In the future, real-time interactive formats created by model developers may overturn the traditional language of AI-animated microdramas.

Zhao Liang: When will real-time interactive multimodal models arrive? Chris, do you have an answer?

Su Guoli: You have asked the right person. Sand.ai chose the autoregressive path precisely with this in mind. Diffusion video generation usually requires waiting until every frame is complete before viewing, while autoregressive generation can stream continuously, generating as it plays. As long as the granularity is reduced far enough—our Magi-1 release in April brought it down to 1 second, and it can eventually reach 1 frame—real-time interaction becomes possible.

Sun Weizhe: Let me respond as well. I divide creators into two categories:

Super-creators: I define them as people at the leading edge. Like fashion bloggers, they connect capabilities across LLM+TTS+ and video models, build workflows, and produce creative content, playing a strong educational and leadership role. We provide them with model and engineering capabilities, and we organize competitions—including the Shaoxing Cultural Tourism Lu You 900th Anniversary Challenge, iQIYI and Paiwo AI's Wonderful Storytelling contest, and the global AI For Good Challenge—and participate in events such as the WAKA Awards, giving creators a stage.

Here is one example. One of our users is an artist living in Florida in the United States. After suffering a stroke in 2023, he tried 17 AI video applications before choosing PixVerse. He later became one of our creators, founded his own art company as a super-individual, and launched an independent product that connects to the PixVerse API for creation. He has received commissions from around the world. That showed me that our product delivers real value to living, breathing people. This is Tech for Good.

Ordinary creators: Tudou's slogan years ago was "Everyone is the director of their own life," but the technology could not deliver on it at the time. Now, with the Agent capabilities of Sora 2 and PixVerse, users can upload one image or enter a simple Prompt to generate an end-to-end story with sound, visuals, subtitles, and a poster. We hope to lower the barrier so that everyone can fully express themselves.

Originally published by Unique Research on Unique Research Substack on November 25, 2025. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research