跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

Why Is 3D Generation a “Source of Pride for China”? Inside the Logic of Monetizing AI Overseas

Original · Unique Research · 2025-11-26

Historical edition note: This complete English edition preserves Unique Research’s reporting and the participants’ views as published on November 26, 2025. References to “today,” “this year,” “next year,” model capabilities and future market outcomes belong to that historical context, not a current update. Company performance, superlatives and forecasts remain attributed to the original speakers.

Editorial clarification: The source’s SIGGRAPH author proportion, tipping statistic, company metrics, valuations and exchange-rate example are retained as historical speaker claims, not independently verified measurements. Cao Renyi is introduced as VP in the opening and as co-founder/product director in his own remarks; both descriptions are preserved. The original’s “Screen Pro” appears to refer to ScreenSpot-Pro, the name used in Google’s evaluation documentation. The Rodin identification remains the original source’s speculation.

Text visible on the original stage backdrop: “3D and Gaming: AI-Driven Industrialized Production and the Boundaries of Creativity.” The event is the 2025 Annual AI Creators Summit and CHINA AI 100 & AI CREATORS 100 Annual Selection, under the UNIQUE BLOOM theme “Pioneering intelligence — the era of the individual,” for “Founder, Creator, Builder & Influencer.” The displayed speaker names and affiliations are retained in the introduction and discussion below.

If you had to summarize the past two years of AI in one sentence, it would be this: AI turned text into content.

But in this discussion about AI applications and 3D generation technology at the 2025 Unique Bloom in Beijing, what we heard sounded more like a preview of something else: the next wave will not be about making content look more real, but about making the world itself generatable, editable, and tradable.

The conversation was moderated by Wang Sheng, a partner at Inno Angel Fund, and brought together four players from very different arenas: Zhao Lu, head of Collov China, who is putting spatial intelligence to the test in home furnishing and the broader real-estate sector; Wang Bihao, co-founder of Xinying Suixing/Doudou AI, who puts multimodal Agents inside game screens to create AI companions that can fight monsters alongside you and understand your emotions; Cao Renyi, VP of LibAI Lab/Cutout.pro, who has carved out a path in the global image and 3D generation market with API delivery in 600ms; and Li Yingli, VAST's chief external partnership representative, whose product Tripo is one of the most representative Chinese forces among today's global 3D large models.

You will find that they have little in common: they operate in different sectors, take different product forms, and follow different commercialization paths.

Yet they are all answering the same question: once AI begins to understand spaces, screens, and objects, where does the real moat lie? In a more dazzling technological paradigm or deeper real-world data? In a larger foundation model or better interaction and delivery?

The people at this table are bringing the answer into view, piece by piece.

I. 3D Is Not Simply a Step Up from 2D; It Samples the Real World

Wang Sheng began with a question about the technological paradigm: Collov previously relied on 2D image generation and now works on Spatial AI. Is it continuing to build 3D from 2D, or is it using new technologies such as 3D Gaussian methods to construct world models?

Zhao Lu's answer ran against conventional wisdom: the route has not changed much, but the underlying logic is clearer. The first step is not to build 3D; it is to understand the real world thoroughly.

He offered a concrete example. In home renovation, cabinets are among the most deeply customized and stylistically diverse product categories. Collov therefore trained its model on 97% of the world's cabinet styles. That may sound like laborious and even unglamorous work, but it produces something solid: a model's recognition of the physical world, super-resolution, and annotation ultimately depend on the density and cleanliness of real data.

There is a larger lesson behind this statement:

World models are not generated through smarter methods; they are accumulated through greater patience.

Tesla's vision algorithms, for example, have not been overwhelmed by general-purpose models. The reason is not that Tesla competes more aggressively on technology, but that its data were gathered on real roads. The same logic applies to Spatial Intelligence. Models are indeed becoming more powerful, but what you feed them determines what they can ultimately do for you.

This is a reminder for many entrepreneurs: the more imaginative AI becomes, the more valuable real-world granularity is. Not every company needs to build the so-called next generation of foundation models, but almost every company can ask itself: in my particular corner of the world, do I possess real-world data that no one else can obtain?

II. The Hard Part of an Agent Is Not the Toolchain but the Mind Behind the Screen

When the discussion turned to Xinying Suixing/Doudou AI—also known as Doudou Companion—Wang Sheng moved the topic from physical space to the screen. On the surface, the complexity of a gaming Agent resembles a technological barrage: it must watch video, listen to speech, understand the environment, remember player preferences, and follow task chains. Yet Wang Bihao began instead with the release of Gemini 3.

He said Gemini 3 Pro's screen-understanding score, Screen Pro, jumped from 11 to 70+, meaning that for the first time a model's multimodal understanding had moved significantly closer to human ability. More importantly, he observed that Gemini's improvement was not merely perceptual but involved learning and reasoning. Give it a wireframe and it can directly produce an aesthetically polished, deliverable front-end project.

That led him to make a weighty observation:

Use Cases are discovered, not manufactured.

In other words, an Agent's value does not lie in stacking up a Workflow. It lies in letting the model judge, within a specific situation, which tools to use and how to help you achieve your goal. The team has even shifted its internal language from Workflow to Reasoning because it wants the model to act in a game world more like someone who understands you than someone who merely knows how to operate tools.

This points to two sources of defensibility:

First, fine-tuning on specialized gaming data, so the model understands game screens rather than merely the real world.

Second, context and interaction design. A general-purpose model may understand cabinets but not the way a designer needs to interact with them. It may recognize a game screen but not understand whether a player's exclamation—“I cannot believe this”—is anger at a teammate or self-deprecating humor.

Put simply, Agents have no reason to fear stronger large models. The danger is failing to build the layer of product intelligence that turns raw strength into usability.

Many people become excited when discussing tool orchestration and long-chain reasoning for Agents, but the real determinant of victory is often more basic: have you encoded users' emotions, habits, and rhythms into the system?

III. The API Business Is Not About How Much a Model Can Do, but Whether It Can Finish in 600ms

When it was Cutout.pro's turn, Wang Sheng asked directly: in the fiercely competitive global Top50 image market, why have you continued to win?

Cao Renyi's answer contained no mysticism, only the language of someone doing business:

First-mover advantage matters, of course, but three things truly make an API business work—scenarios, speed, and Know-how.

Scenario differentiation is easy to understand. Within 3D generation, interior design demands centimeter-level precision, while games demand explosive imagination. Give both groups the same API and both will reject it.

Speed is even more concrete. Cao cited POD—print-on-demand—e-commerce: someone buying a T-shirt inserts a portrait and wants to see the result immediately. Make that user wait 10 seconds and they will already have placed an order elsewhere. Cutout.pro's API returns an image in 600 milliseconds. In e-commerce, that is not called optimization; it is the line between life and death.

Know-how, meanwhile, is a kind of judgment that combines the engineer, product manager, and commercial operator:

You must know what technology can and cannot do at its current stage.

Realistic capture of subtle human facial expressions is still imperfect, for example, so do not force a photorealistic route. But cartoon styles are already reliable, so take them as far as they can go.

This is not a compromise. It is treating technological boundaries as commercial boundaries that must be mastered.

It may sound cold, but this is reality:

Once AI enters an industrial value chain, model capability is merely the admission ticket; delivery efficiency is the moat.

Do not underestimate those few hundred milliseconds. Behind them lie conversion rates, repeat-purchase rates, average order value, and even whether you survive until the next funding round.

IV. Why Has 3D Generation Become a Source of Pride for China? Because China Has Talent, Data, and Industrial Pressure at the Same Time

When the topic reached VAST / Tripo, Wang Sheng asked a question that sounded admiring but was actually quite sharp: in the competition between Chinese and American models, the United States leads in text and multimodality, but Chinese companies appear to have the upper hand in 3D generation foundation models. Why?

Li Yingli divided her answer into three layers:

an advantage in talent, an advantage in data, and the pull of industry.

More than 60% of authors at SIGGRAPH are of Chinese heritage, which effectively means that the main talent pool at the intersection of 3D+AI lies in the Chinese-speaking world.

China is also an enormous repository of 3D industrial scenarios: gaming, e-commerce, industrial design, 3D printing, embodied intelligence, and more. These value chains generate high-precision model data every day.

More importantly, industrial demand is urgent. How urgent? Urgent enough to force rapid deployment of technology, which would otherwise be overwhelmed by the next wave of demand.

But she also stated a point worth pausing over for a couple of seconds: the technology is in China, while the applications are overseas.

Foreign users account for 85% of Tripo's users. You can sense a little frustration in that figure, but also see a trend:

China leads in 3D foundation technology, but the first explosion of spending power and creative demand may still come from elsewhere.

What does that mean?

It means 3D generation is likely to become another essential route for exporting technology: models iterate in China while markets deliver returns overseas.

As several guests noted, subscription systems, infrastructure, payment habits, and exchange-rate advantages overseas mean that AI applications will probably first make substantial money abroad and then return to deepen the industry at home.

V. The Story of 3D Printing Matters Because It Turns “Everyone Is a Creator” into a Hardware Reality for the First Time

The conclusion of this conversation moved 3D generation from a software vision into hardware reality.

VAST is working with many leading 3D printer makers: Bambu Lab, Anycubic, Creality, and others. Even industrial giants such as Stratasys have become strategic partners.

Li Yingli's judgment was direct:

The reason 3D printers failed to become widespread was not inadequate machines, but a barrier that made modeling too difficult.

Once AI removes that barrier, the printer market will move from a toy for a small minority to a household tool.

You can compare it to the adoption of paper printers:

If everyone can turn a small object imagined in their mind into a 3D model with one click, the printer will no longer be a geek's toy. It will become the new microwave oven in the home.

This carries major significance for the AI industry:

when AI turns creativity into a repeatable everyday ability, an industry's growth curve can suddenly move from a gentle slope to a near-vertical ascent.

That is because it does not merely make you more efficient. It gives you the right to participate for the first time.

Finally, Put the Four Companies Back into the Same Picture

If I had to compress this conversation into one sentence, I would say:

the next stage of AI applications will not be decided by who has the largest model, but by who understands the details of the world best.

Collov samples the world through specialized real-world data;

Doudou AI accompanies the world through multimodal reasoning;

Cutout.pro delivers the world through speed and scenarios;

Tripo expands the world through its 3D foundation and industrial value chain.

The four companies are standing guard at different gates, but they protect the same city:

the entrance to a new civilization moving from generating content to generating spaces, objects, and experiences.

For those of us watching, working in the industry, or building companies, the three ideas most worth taking away may be:

depth of data, warmth of interaction, and speed of delivery.

Models will keep becoming stronger; that is almost certain.

But whether you can turn that strength into something real users want to use, pay for, and depend on is the question you truly need to answer.

After all, the world has never been changed by models.

It is changed by people who use models correctly.

More Details from the Conversation

I. Guest Introductions

Wang Sheng: I am very pleased to join this major AI application event organized by Mr. Wu Wei and Unique. We are all old friends. Looking at today's guests, I can see that they were carefully selected: they are closely connected to our topic and are all outstanding leaders in their industries. Let us begin by asking everyone to introduce themselves for one minute, and we will then expand on the discussion.

Zhao Lu: Thank you, moderator, and thank you, Mr. Wang. I am Zhao Lu, head of Collov China, where I am also responsible for product commercialization and marketing. We are a Silicon Valley company. Some time ago, many of you may have seen an article by Fei-Fei Li arguing that Spatial Intelligence will define the next decade of artificial intelligence. We began developing in this direction as early as 2021. We both build our own underlying visual large models and develop vertical applications.

Across our global ToC and ToB businesses, we have registered several million users in total. We have completed four funding rounds and serve more than 20,000 SMBs. Our main focus today is using spatial-intelligence technology to serve the broader real-estate, home-furnishing, and home-renovation sectors. Thank you.

Wang Bihao: Thank you, Mr. Wang. I am Wang Bihao from Doudou AI. Our product Doudou Companion is an AI companion for gamers. By sensing the game screen and voice in real time, it provides gameplay guidance and emotional support. Our overseas product, Hakko.AI, has been used by more than 10 million players worldwide, with monthly active users—MAU—also numbering in the millions.

For both products, the principal capability is multimodal perception. It is the same concept as the recently popular Google Gemini 3. I will talk later about some of the lessons we have drawn from this Gemini release.

Cao Renyi: Hello, everyone. I am Cao Renyi, co-founder and product director of LibAI. Our company name may not be familiar to everyone, but some of you may have heard of our product, Cutout.pro. We are the only Chinese company to appear on a16z's all-star list 5 times, and we receive more than 14 million visits per month. We primarily provide APIs for background removal, image enhancement, and the 3D generation related to today's theme. We can discuss more later. Thank you.

Li Yingli: Hello, everyone. I am Li Yingli from VAST, a globally leading company building AI large models for 3D generation. Our product is called Tripo and enjoys a strong international reputation. The Tripo platform has gathered more than 5 million creators, more than 40,000 small and medium-sized customers worldwide, and more than 700 major customers. The number of AI 3D models generated on our platform has exceeded 100 million.

In fact, whether measured by its team, product capabilities, user feedback, or the scale of API usage, VAST's Tripo ranks first worldwide and is an industry leader.

Wang Sheng: Thank you to all four guests. Let me also spend one minute introducing myself. I am Wang Sheng, a partner at Inno Angel Fund. Inno primarily invests in early-stage startups, and we are willing to write entrepreneurs their first check. The amount of a single first-round investment is now approaching RMB 20 million, and we currently manage around RMB 6 billion or more in capital. We have invested in many leading AI robotics and embodied-intelligence projects.

Turning specifically to today's 3D and gaming themes, we were an early investor in VAST. An application recently released for HarmonyOS—editor's note: the source speculates that this may refer to Rodin or a similar 3D application—has ranked first in HarmonyOS downloads over the past month, ahead of WeChat and Douyin. It is a breakout product for 3D reconstruction and was also one of our first-round investments. In gaming, we were an angel investor in Yoozoo Games as well as Microfun (柠檬微趣), which generates several billion yuan in annual revenue.

To summarize the lineup: we have Collov AI, an AI home-design business that can be described as a Spatial Intelligence company; Xinying Suixing, which offers gaming companionship and is a classic Agent application; LibAI/Cutout.pro, a leader in image generation; and Tripo (VAST), a leader in 3D model generation. This is an exceptionally strong lineup, and I am sure everyone here possesses deep expertise.

II. Technology Routes and Product Moats

1. The Technological Paradigm for 3D Generation

Wang Sheng: My first question is for Mr. Zhao of Collov AI. As I recall, you previously used 2D image generation to produce higher-quality, highly controllable, and highly consistent images. Now that you are building Spatial AI, has your technological paradigm changed? Are you continuing to reconstruct 3D from 2D, or will you use new technologies such as 3D Gaussian Splitting to construct it?

Zhao Lu: In fact, our technology route has never changed very much. We need to return to what we just discussed: why we focus on the business system surrounding real estate and home furnishing.

Looking purely at computer vision, many recognition algorithms are used in areas such as production-line inspection. In the United States, companies dedicated to inspecting and repairing roofs have also become very large. We believe that doing spatial intelligence well first requires mastering the underlying algorithms for image recognition, super-resolution, and annotation. In other words, it needs an exceptionally clear data foundation for the physical world itself.

That is why we trained on 97% of the world's cabinet styles: cabinets are the most highly customized part of a home-renovation system. We go deeply into a specialized field, applying experience and data from the real physical world before moving into the virtual realm and crossing from one to the other. This is our technical path: first enable the real physical world and collect its data well.

Wang Sheng: That is very interesting. A major school of thought in building a World Model is to construct 3D from 2D. But with so many companies working on image generation—including LibAI, Google, and Midjourney—will these foundation models not squeeze you out?

Zhao Lu: It depends on which specialized datasets each company ultimately cultivates. Consider Tesla: if it built only the most basic automotive vision model, NVIDIA should theoretically have crushed it long ago. Yet Tesla's visual algorithms still lead because of its data. The same applies to us; we focus on accumulating data in specialized industries.

2. The Complexity and Multimodality of Agents

Wang Sheng: My second question is for Bihao. My understanding is that you are building a true Agent, currently the hottest direction but also a very complex one. You need to understand game video, speech, and the environment, while also learning player preferences, emotions, game-task Memory, and more. How complex is this? What happens if Tencent comes in to compete with you?

Wang Bihao: You cannot discuss Agents without large models. I would like to begin with Gemini 3. In this release, Google emphasized Multi-modality and three levels of capability: Learn, Build, and Plan, while DeepMind's Hassabis discussed Benchmarking as a discipline.

The improvement is exceptionally significant. Gemini 3 Pro's Screen Pro score—its screen-understanding capability—rose from 11 to more than 70. This means the model's understanding of screens and multimodal information can now rival human ability.

Our testing found that it is particularly strong at generating front-end projects. Take a WeChat mini-program poster generator, for example. Compared with other models, Gemini 3 has exceptionally good aesthetic judgment. Give it a simple wireframe and it produces something that can be delivered commercially. This shows that Google treats the capability as learning, not merely perception.

That is exactly where we can enter. Use Cases are discovered, not manufactured. We no longer call it a Workflow; we call it Reasoning. The large model decides which tool to use for which task.

Our moats are:

Data tuning: we have a model system tuned with a large volume of specialized gaming images. This enables the large model to understand a game screen rather than interpret it as a general model interprets the real world.

Context and Interaction: your interaction design and contextual understanding must outperform a general-purpose large model. As Mr. Zhao just said, a general model can generate cabinets, but it cannot be used without interaction optimized for designers. The interaction must be better than what the large model offers.

Wang Sheng: I strongly agree. A friend of mine, Yang Yue—described as China's leading figure in AI music—has also said that however strong a foundation model becomes, competition ultimately comes down to aesthetics. Why is Midjourney so good? Because its aesthetic standard is high.

3. Differentiating APIs through Vertical Scenarios

Wang Sheng: My third question is for Mr. Renyi of LibAI. Cutout.pro has long ranked in the global Top 50 for image products, an extraordinarily competitive battlefield. Why are you so successful and able to compete even more aggressively than others?

Cao Renyi: We began going global in 2019, so we have a first-mover advantage.

First, scenario differentiation. We provide APIs and go deep into different industries. In the 3D field, for example, interior design and gaming have completely different requirements. Interiors demand precision—a ceiling accurate to within a few centimeters and one-to-one reproduction—while gaming demands imagination and visual impact. As large-model technology advances, vertical industries have greater opportunities to penetrate distinct scenarios.

Second, speed and conversion. Consider POD—print-on-demand—businesses in e-commerce, where users want to print a portrait on a T-shirt. If a large-language-model API takes 10 or 20 seconds, the customer has already left. Our API can return an image within 600 milliseconds, which has an enormous effect on e-commerce conversion rates.

Third, Know-how. Consider Fashion Try-on and Furniture Try-on. Creators or short dramas may need to keep a character ID unchanged while exercising imagination. In video, current capabilities cannot yet capture subtle human facial expressions perfectly, but they can deliver a Disney-style cartoon look extremely well. We must understand what technology can and cannot do and maximize value within the limits of its capabilities.

4. Why Is 3D Generation a Source of Pride for China? A Question for VAST

Wang Sheng: In competition between Chinese and American large models, the United States leads in text and multimodality. Yet in 3D generation foundation models, China is far ahead in technological leadership, user scale, and the number of startups, with VAST as one representative. Why has 3D generation become a source of pride for China?

Li Yingli: AI 3D sits at the intersection of artificial intelligence and computer graphics.

Talent advantage: anyone who has attended SIGGRAPH, the premier computer-graphics conference, knows that more than 60% of authors on global AI projects are of Chinese heritage. Our company has brought together the world's best young scientists from this group.

Data advantage: China has exceptionally rich 3D-related industries—including gaming, e-commerce, industrial design, and 3D printing—and is a data powerhouse. The diversity and flexibility with which data can be obtained gives us the world's largest collection of high-precision 3D model data.

Industrial pull: we are eager to empower industry, while demand from Chinese industries such as embodied intelligence and 3D printing has in turn driven our rapid development. Objectively speaking, however, although our technology leads, 85% of users on our platform are still foreigners. In other words, the technology is in China, while the applications are overseas.

III. Commercialization and Markets: How AI Applications Make Money

1. Collov's Overseas and Domestic Strategies

Wang Sheng: Collov was founded in Silicon Valley and has developed well in North America. Mr. Zhao, as head of China, do you help clients expand overseas in addition to providing AI capabilities? What is the situation in China?

Zhao Lu: Collov's DNA comes from Stanford's AI Lab. Before Fei-Fei Li founded her company, we had already open-sourced a project related to the World Model.

Our overseas business is now growing rapidly. Beyond North America, Brazil, the Middle East, and Central Asia are also growing particularly quickly, at roughly 30% to 50% per month. On Google AdWords, we dominate several core keywords.

Our domestic strategy resembles Tesla's entry into China:

Compliance first: we obtained the relevant certifications relatively early.

Win the leaders: first work with governments, major industrial clusters, and publicly listed companies to create a lighthouse effect, then expand down the market to mid-sized merchants and studios.

Different models: overseas, everything is based on pure License/SaaS models; in China, virtual and physical delivery are combined through project work. For example, with a Beijing decorating company backed by state-owned capital, we implemented a flagship store covering more than 10,000 square meters. The China business needs to be heavier, unlike pure SaaS delivery overseas.

Wang Sheng: Do you worry that money is still easier to make from foreigners?

Zhao Lu: Indeed. People overseas are accustomed to paying. In China, even an outstanding product often needs a matrix of lightweight applications to validate the market.

2. Xinying Suixing's Growth and Monetization

Wang Sheng: You gained several million users within one week of launch. How did you do it, and how are you thinking about commercialization?

Wang Bihao: On growth:

Content: the core of content is sincerity and leverage—the ability to use content to activate a KOL or Influencer, then expand its reach.

Cold Start: create momentum around a topic so that people who spread it believe it is worth sharing.

We cannot spend enormous sums on paid acquisition each month as ByteDance does for Doubao, so we rely more on leverage.

On commercialization:

Yesterday, I spoke with an American journalist and told him, “You need to understand that China's ToB ecosystem is different.” In the United States, you can find many small businesses willing to pay for SaaS. In China, small businesses find it difficult to access high-end AI services and have little willingness to pay. Chinese consumer users also currently have limited spending power. Americans, however, have a Tipping Culture, and 87% are willing to tip US$30 every month.

That is why our current priority is the US market and subscriptions.

3. LibAI's Payment Challenge

Wang Sheng: Chinese users genuinely dislike paying, whether businesses or consumers. Mr. Renyi, can you actually collect the money?

Cao Renyi: The overwhelming majority of our revenue comes from overseas.

Commercial environment: overseas customers are accustomed to using external services because labor is expensive and transaction costs are lower than management costs.

Infrastructure: overseas markets have mature Subscription systems that can automatically charge users each month through MRR/ARR models. In China, apart from giants such as WeChat and Alipay, ordinary companies find it difficult to ask users for automatic monthly payments.

Exchange rates: US$1 equals RMB 7, and there is also a difference in purchasing power.

Wang Sheng: A quick survey of the room—raise your hand if you have paid for an AI product. There are quite a few of you, but you are far from typical. Making US dollars through overseas expansion is indeed the current consensus.

4. VAST's New Commercial Exploration: 3D Printing

Wang Sheng: I hear that VAST is working with all the leading domestic 3D printer makers. Does this appear to be a new source of growth?

Li Yingli: Yes. The household 3D printer market is certain to experience explosive growth in the future.

In the past, there were only a little over 10 million 3D-printer users worldwide. The bottleneck was the 3D model itself, an elite art that ordinary people could not create. Now, with AI, complete beginners can generate models and everyone can become a creator. That will propel the printer market from millions of users to billions.

We have cultivated China's 3D printing industry deeply. Bambu Lab, Anycubic, Creality, and ELEGOO are all our customers. The global industrial-printer giant Stratasys is also our strategic partner. In this segment, our market share is extremely high and basically approaches 100%.

IV. Conclusion

Wang Sheng: Excellent. Like the paper printers of the past, 3D printers will enter the home in the future. Companies such as Bambu Lab already have valuations of US$10 billion. I hope everyone will be able to use them as costs fall. I also wish every guest here great commercial success, especially success in making money in China.

Originally published by Unique Research on Unique Research Substack on November 26, 2025. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research