Original · Unique Research · 2026-03-23 · Shanghai
Editor's note: This is the complete English rendition of Unique Research's March 23, 2026 interview-based article. First-person commentary belongs to the original author. The headline says Vozo topped Product Hunt three times, while the body identifies the November 2024 result as second place in a monthly ranking; both source statements are retained, and the discrepancy has not been independently resolved. Most of the source names Zhou Changyin, but one introductory sentence says Zhou Chang; that variation is preserved. Shuode AI is a romanized rendering, not a verified official English brand name. Career credentials, rankings, user counts, ARR, funding and positive cash flow are source claims rather than independently audited findings; ARR is not cumulative revenue, and positive cash flow is not a claim of accounting profitability. The market estimate written as 100B has no currency specified in the source. Market-growth expectations and future killer apps remain the interviewee's predictions. Back Translation is presented as an additional check, not proof that a translation is correct, and the source's explicit limits on trusting AI remain in the body.
Unique Awards · Guest Interview
He Worked on Gcam and Glass at Google X. Now He Wants to Take One Video Directly to the Whole World
Zhou Changyin: True video globalization is not about translating subtitles into a foreign language. It is about making content enter every market as if it were native to that market.
Interviewee: Zhou Changyin (Founder of Vozo AI)
Interview date: March 2026
If you still understand "taking video overseas" to mean finding someone to translate the subtitles, adding a foreign-language audio track, and editing a few localized versions,
you may be underestimating how this is about to change.
The next generation of AI video platforms is no longer trying to solve "how to translate a video." It is trying to turn one piece of content directly into native content for multiple language markets.
That was one of my strongest impressions after speaking with Vozo AI founder Zhou Chang.
Not Everyone Building AI Video Truly Understands Video
Zhou Changyin's background is relatively unusual among today's AI entrepreneurs.
He earned his bachelor's and master's degrees at Fudan University and a doctorate at Columbia University, specializing in computational photography. He was an early scientist at Google X and worked on Gcam and Glass. He then became a serial entrepreneur, moving from VR in 2015 to Shuode AI in 2021 and Vozo AI in 2024.
On the surface, that path spans several fields: computational photography, VR, video creation, and AI video globalization. Viewed together, however, they reveal a remarkably consistent thread:
His work has always been about "making video a stronger, better, and more accessible form of expression."
The problem differed at each stage. In the computational-photography era, the focus was how to capture better video. In the VR phase, it became how to create "immersive video." With today's Vozo, the direction has advanced another step:
How to make it easy for everyone to express themselves through video, across the boundaries of language and markets.
That is why Zhou Changyin places special emphasis on one principle:
Technology can be very complex, but the user experience must be extremely simple.
Behind Three Product Hunt Launch Campaigns
📈 Vozo topped Product Hunt upon its July 2024 launch, reached US$1 million in ARR within 6 months, and has more than 6 million users worldwide.
🏆 In November 2024, its second product upgrade placed second in Product Hunt's monthly ranking.
🥇 On March 10, 2026, it launched Visual Translate and topped Product Hunt again.
But Zhou Changyin's own description of reaching the top the first time is restrained. He says expectations were not particularly high: Vozo Rewrite had just come out, and the team simply thought the capability was interesting and might go viral. When Product Hunt CEO Rajiv personally voted for and shared it, the greatest significance for the team was not "the traffic is here," but recognition of the product experience by someone who truly understood products.
By the third launch in 2026, Zhou Changyin's expectations were much firmer. The reason was straightforward: the team already had a sufficiently large user base and knew users really needed the ability to "translate text inside video frames."
After Sora Took Off, Why Did They Become Even More Committed to the Application Layer?
When Sora launched, many entrepreneurs felt threatened. Zhou Changyin's judgment, however, was clear and somewhat counterintuitive.
He says that before Sora launched, the team had already spent almost a year developing a video-generation model called HiveNet. So his first reaction to Sora was not "the sky is falling," but that the technical approach itself was not especially innovative, yet was very sound and handled many details well, producing outstanding results.
In his view, video models will eventually resemble today's LLM landscape: several leading providers supplying foundational capabilities like electricity and water. What determines the depth of the market is not "whether utilities exist," but how many application-layer systems emerge on top of them for different scenarios, needs, and users.
Do not build something a train will run over. Build the rich network of roads along the railway instead.
Real PMF: Not Whether Your Idea Seems Reasonable, but Whether Someone Will Pay for a Crude Version
Zhou Changyin says there are two very direct ways to ensure you start with genuine demand:
First, find a group of real people with needs they are genuinely willing to pay to meet.
Second, build a very crude product and get them to actually pay.
The wording is blunt, but it rings true. One of the easiest illusions in AI entrepreneurship is that an idea makes logical sense, is technically feasible, and appears to have a market. But "reasonable" does not mean "real."
AI's Biggest Problem Is Not Intelligence, but Whether Users Dare to Trust It
One representative Vozo feature is called Back Translation.
Zhou Changyin's judgment is straightforward: one of AI's biggest problems today is that users do not trust it. More precisely, AI often does not deserve complete trust in the first place.
Back Translation is therefore a pragmatic fallback when there is no better option. It is not the ultimate answer, but a temporary patch for trust. It gives users at least one additional layer of checking, so they know what AI has actually translated for them.
After Two Ventures, His Biggest Change Is Not Better Technical Skills, but a Greater Ability to Let Go of Ego
Zhou Changyin says product management may be the most interesting job of this era, precisely because it sits between "identifying needs" and "technical possibility."
He consequently developed an important criterion: try to find the position where "the whole industry is missing this one link." If you fill the most critical, most absent link in the entire pipeline, the timing risk becomes much smaller.
This also reflects another lesson from his own experience: choose the beginning or the completion stage of a pipeline where possible, rather than getting stuck in the middle.
A 40-Person Team and Positive Cash Flow: What Did This Company Get Right?
In today's environment, seeing an AI startup reach positive cash flow with a 40-person team and US$6 million to US$7 million in funding inevitably prompts the question: how?
Zhou Changyin's answer is unflashy, but forceful:
When you discover you are wrong, change course immediately. Burning money cannot rescue the wrong product.
AI Video's Real Explosion Has Not Yet Begun
Many people already describe competition in AI video as white-hot. Zhou Changyin's judgment is the opposite: penetration is still very low, and "there aren't many decent products" today.
He even thinks people's imagination of the Video Production market falls far short. Many currently put it at 100B; he considers that a serious underestimate, with potential for growth of 10 to 100 times in the future.
Vozo's Real Bet Is Not on Translation, but on Video Becoming the Foundational Language of the Next Generation of Global Expression
Toward the end of the interview, Zhou Changyin mentioned a detail worth considering.
Vozo started with the US market because US SaaS users were more representative globally. Over the past year, however, the team noticed organic growth among users in China too, and decided to formally support them by completing the Chinese-language experience and service offering.
On the surface, this looks like simple market expansion. Underneath, it reveals another judgment:
They do not see themselves as an "overseas-expansion tool," but as a platform for video expression across the world's different language markets.
China, the United States, and Europe are only starting points in different regions. The larger proposition is that content will increasingly become video, and video will become increasingly global.
At that point, who is doing "translation" will no longer matter. What matters is who is redefining "native content."
So if I had to define Vozo in one sentence, I would rather put it this way:
Vozo is not helping you translate a video for overseas use.
It is helping you take that video directly into another linguistic world.
And that may be the real opportunity for AI video applications.
Selected Q&A
Q1: What exactly is Vozo?
Zhou Changyin: Vozo is a platform that uses AI to "globalize a video in one click"—automatically handling dubbing, lip synchronization, subtitles, and translation of on-screen text, so a single video can enter every language market like native content.
Q2: What were your expectations when you first topped Product Hunt?
Zhou Changyin: They were not particularly high. We were just getting started and thought Vozo Rewrite was an interesting capability with a chance to go viral. When Rajiv personally voted for and shared it, what mattered most to us was recognition of the product experience.
Q3: Why were you more confident about the third launch?
Zhou Changyin: Because we already had a relatively large user base and knew people genuinely needed "visual translation." When we launched Visual Translate in March 2026, we expected a positive response, and the result largely matched our expectations.
Q4: What is the consistent thread in your product philosophy from Google X to Vozo?
Zhou Changyin: Technology can be very complex, but the user experience must be very simple and intuitive. Technology should be developed around user needs and experience, and delivered completely. Users should not have to compensate for its shortcomings.
Q5: Why focus on vertical scenarios rather than building a general-purpose visual model?
Zhou Changyin: Foundation models will become like electricity and water, supplied by a few leading players. They are suited to major companies, preferably cloud providers. The places where real profits can be made are the computing infrastructure and the application layer that captures user demand. We prefer solving real market problems.
Q6: Many Vozo capabilities look more like engineering optimization than model innovation. Why does that matter?
Zhou Changyin: Once models become too standardized, they easily become commodities and struggle to create barriers to competition. A great deal of seemingly mundane data, engineering, and systems work is actually essential to an excellent user experience—and is becoming a scarce technical barrier.
Q7: How do you judge whether a need is real rather than imagined?
Zhou Changyin: The most direct method is to find a group of real people with needs they are genuinely willing to pay to meet, then build a very crude product and get them to actually pay. A business does not exist simply because it is "reasonable." Many factors you have overlooked could kill it outright.
Q8: Why build Back Translation?
Zhou Changyin: Because trust in AI is a major problem. Back Translation is a fallback when there is no better option; at least it helps users add a layer of checking. Especially in video translation, whether users can trust the result is critical.
Q9: Why did it take so long to address the need to "translate text inside video frames"?
Zhou Changyin: Because it is very difficult. Text in video frames can take many forms, have complex backgrounds, and move. The need had always existed, but only in Q4 2025 did we judge that the relevant visual-understanding and generation technologies had reached an acceptable threshold. We then decided to build it, eventually releasing it in March 2026.
Q10: Why did you abandon a similar direction in 2022, then return to it?
Zhou Changyin: At the time, we took representative videos and manually translated them using the best available technology. We found many aspects below standard—for example, the TTS after Voice Clone was not good enough. A year later, Voice Clone technology had improved considerably, and our accumulated capabilities in video understanding and AI editing had strengthened, making it possible to turn this into a product-level capability.
Q11: What was the biggest lesson from two ventures?
Zhou Changyin: Founders with strong R&D backgrounds can easily get too far ahead. In a startup, try to find the position where "the entire industry is missing this one link," rather than placing yourself in the middle, waiting for upstream and downstream capabilities to mature together.
Q12: How do you interpret "sunk costs should not influence major decisions"?
Zhou Changyin: Abandoning a product line certainly takes courage. You have to admit an error of judgment and take responsibility. But burning money cannot rescue the wrong product. When you discover you are wrong, changing course immediately is usually the better choice.
Q13: Why do you say product management may be the most interesting job of this era?
Zhou Changyin: Because product managers stand between identifying needs and technical possibility. That is the most challenging and interesting position. The hardest part is genuinely moving from "what can I do?" to "what do users need?"
Q14: How do you see the opportunities in AI video over the next three years?
Zhou Changyin: Market penetration is still very low today, and there are not many genuinely decent products. The key variables ahead are continued improvement in model capabilities and the emergence of 3 to 5 killer apps. Once those two things come together, the market will become far larger than people currently imagine.
Q15: If you were starting an AI product from 0 today, how would you choose a direction?
Zhou Changyin: Choose something you enjoy or a market you know, then do it again using the latest AI capabilities.