跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

Zero AI Hardware Has Hit a Million Units — Why Are These Founders More Excited Than Ever?

Original · Unique Research · 2026-08-05

Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening essay, four insight sections, closing reflection, and the full roundtable transcript. All named companies, products, and people are preserved. Product sales figures, funding amounts, star counts, and customer counts are speaker claims, not independently verified findings.

AI Industry Observer

"No million-unit hit product may be precisely because previous AI hardware was selling features, not operating data loops."

The AI Pin flopped. The Rabbit R1 flopped. Lookie and Nuna gather dust at high rates. Even Plaud is still the voice recorder + AI summary form from late 2023.

A stark reality: to date, no proactive multimodal agent hardware has sold a million units worldwide. The concept of AI hardware is now hyped into something fantastical, but the products that have landed are all very grounded.

These were the words of Star Dome Ark CEO Duan Ran (段然) at a roundtable. By all logic, this number should have made the hardware founders in the room pessimistic. But at the same table — the force-control robotic arm maker, the education robot maker, the brain-computer interface maker, the home companion hardware maker — not a single person was pessimistic.

The reason is simple: none of them are betting on the model. They're betting on data.

Factories Are Not the Promised Land of Embodied Intelligence

Start with a counterintuitive judgment.

The hottest story in embodied intelligence right now is: large models bring generalization, and robots can suddenly do everything. But Flexiv's Hu Xiaoping (胡晓平) poured cold water on the spot: most factories want precision and stability, which is fundamentally mismatched with large models' generalization goal of "doing both A and B."

Flexiv spun out of Stanford Robotics Lab, spending ten years obsessed with force control to make robotic arms do dexterous operations like human hands. Hu gave a vivid example: when a person plugs in a USB, they don't know the exact position of the port, but they get it in every time. The perception of force, the touch technique — these are etched into human bones, but putting them on a robot is extraordinarily difficult.

Traditional industrial robotic arms follow a precision-position route, which means they can only work in fixed scenarios and are helpless in open environments. That's why so many labor-intensive steps in factories remain beyond robots' reach.

"His conclusion was direct: over the next three years, factories are the testing ground, not the volume-scaling ground, for large models and embodied intelligence. The first to generate scaled revenue will be robotic arms that have operational capability pushed to the extreme, not chatty robot bodies. Yield and reliability are just the entry ticket; the only metric that drives procurement is payback period."

Context Is All You Need

Duan Ran was the most outspoken person at the table.

He does AI wearables — glasses, pendants, rings all in the pipeline. Outsiders see this as resource dispersion, but his answer is a very effective demystification: there is no grand narrative of "finding the best carrier." It's simply that we don't know which product users will like, so we launch several — as long as each can recoup costs, it's fine.

Why does he dare do this? Because of Shenzhen. Chips uniformly use the few from BES, JieLi, Actions, Allwinner; the sensor supply chain is also from Goertek and AAC. Over the next five years, an AI-native hardware product like a computer — where the sensor supply chain doesn't change, and adding one SKU costs not 80% or 100% more but only about 20% — will appear. The first consideration in business is cost control. That grounded.

Demystified as it was, his judgment is unambiguous. He believes that over the next three years, an AI-native hardware product like a computer, owned by over 70% of young people, will emerge. The specific form? He won't say: "If I knew, I'd already be ringing the IPO bell."

"But what he dares say is the direction: the core of next-generation AI hardware is context. When base model capability is the same, whoever holds users' all-day context data will have service capability over 40% higher than others, and users will never leave you. Among all product forms, only wearable hardware can achieve this."

What about giants circling? He offered two paths. One is to learn from Even Realities, targeting European business professionals — a high-net-worth segment that just raised 1 billion RMB two months ago. The other is more direct: hold onto a giant's coattails. A new product will be launched at Huawei's HC Full Connect Conference in September — welcome everyone to try it then.

The one sentence from the entire roundtable I most want to remember is also his: "There are no narrow tracks, only your narrow field of vision." No matter how narrow the category, going global from day one gives you a base large enough to support a company.

Hardware Is Not a Moat — and That's Coming From a Hardware CEO

WhalesBot's Jia Xindong (贾昕东 Slade) raised an even more counterintuitive point.

WhalesBot makes youth education robots, with 400+ products sold in 81 countries and seven consecutive years of profitability — surviving the pandemic and the double reduction policy. By all logic, hardware should be its lifeline. But Jia Xindong says: hardware is indeed an advantage overseas, but content is the deeper, longer-term moat.

The logic is elegant. Precisely because China has a one-of-a-kind supply chain in the Pearl River Delta and Yangtze River Delta, anyone can make hardware cheap — which means it doesn't constitute a barrier. Doing only hardware overseas makes it hard to break out.

What's truly hard-core is content and data. A self-developed AI platform filters all content unsuitable for children, with values that don't chase scores but cultivate AI thinking; the self-run ENJOY AI global youth AI robotics innovation competition has run for seven years, with over 150,000 teenagers participating annually from around the world; over 30,000 schools and 10,000 institutions domestically use their solutions.

Next, targeting the consumer home market, he says openly that he's benchmarking Imoo (小天才): on the parent side, hitting safety anxiety; on the child side, hitting social and entertainment — capturing both ends. And in the AI era, seven years of accumulated youth data will be this company's most valuable asset.

Two "Dumb Approaches," the Same Data Path

There were two more — one drilling into homes, one drilling into the human brain. Their approaches look the slowest, but the logic is identical.

Fenghuolun Yingtu's Kong Weigang (孔维刚) does proactive interaction: most smart hardware still does nothing until you call it; they want devices that perceive context and proactively serve. His breakdown is hard-core: phones are designed for communication — with the screen off, how do you continuously perceive? Phones can't be fixed in one scene to solve the angle problem over time; the data source is inherently flawed. The primary carrier for proactive interaction must be a new form of agent hardware.

But selling hardware isn't the goal. What he truly values is the loop behind the hardware: real demand drives high-frequency use, high-frequency use produces real scenario data, data feeds back into the proactive interaction brain, and the brain makes the hardware smarter.

Shenwu Tech's Li Yang (李杨) does brain-computer interfaces. The commercialization path is equally "dumb": don't touch consumers first; sell to research. Tsinghua, Peking University, and other universities are all building BCI disciplines, using their equipment for experiments. Because brain-computer large models are the industry's inevitable path, whoever first occupies data quality and builds a data platform will occupy the upstream of the industry.

Closing

Five people, five paths. Robotic arms, glasses, education robots, companion hardware, EEG headbands — the product forms have nothing in common.

"But take each person's words apart, and underneath is the same judgment: model capability will converge, supply chains will be flattened, and the only thing that can't be flattened is who is closer to real-world data. No million-unit hit product may be precisely because previous AI hardware was selling features, not operating data loops. No one knows what the next truly breakout AI hardware will look like. But its owner will most likely be the person who has been accumulating data from day one."

More Conversation Details

Speakers

Fenghuolun Yingtu Founder & CEO Kong Weigang (孔维刚)

Star Dome Ark CEO Duan Ran (段然)

Flexiv Group VP Hu Xiaoping (胡晓平)

Shanghai WhalesBot AI Robotics Overseas 2C Market Lead Jia Xindong Slade (贾昕东)

Shenwu Tech Partner & VP Li Yang (李杨)

Host

Unique Research Shenzhen Lead Kang Zhengzhong (康正中)

Kang Zhengzhong: Before formal discussion, please each introduce yourself and your business in about one minute. Mr. Kong first?

Kong Weigang: Hello everyone, I'm Kong Weigang, founder & CEO of Fenghuolun Yingtu. We build a commercial closed loop based on a proactive interaction brain, real-scene interaction data, and home companion agent hardware. The core is to drive agent hardware and robot bodies from instruction perception to proactive interaction through underlying technology.

Hu Xiaoping: Good morning everyone, I'm Hu Xiaoping from Flexiv. Flexiv was founded in 2016, spun out of Stanford Robotics Lab. We mainly solve robotic manipulation capability problems. As agents land in the physical world in the future, manipulation capability empowerment is critical. The core logic is making robotic arms closer to human hands, breaking through technical innovation in the force-control dimension, so robots complete dexterous, flexible tasks like human hands.

Jia Xindong: Hello everyone, I'm Jia Xindong, head of overseas marketing at Shanghai WhalesBot AI. We focus on AI + education, founded in 2017, with very successful commercialization. The company has 400+ robots, covering humanoid, flying, modular, and companion mobile types, with partnerships in 81 countries worldwide.

Li Yang: I'm Li Yang from Shenwu Tech. We do brain-computer interfaces, backed by Tsinghua's Neural Engineering Network Lab, applying technology to real life. We hope BCI makes human brains healthier and stronger, connecting you and me, connecting everything.

Factories Are Not the Promised Land of Embodied Intelligence

Kang Zhengzhong: First perspective — the convergence of agents with new carriers, how do you view AI and hardware fusion innovation? Mr. Hu first. Flexiv is currently in the hottest embodied intelligence track. Everyone is chasing humanoid robot bodies, but Flexiv chose to put human-like capability into seven-axis robotic arms, turning uncertainty into adaptability through force control, mechanics, vision, and AI. From the embodied intelligence or robotics perspective, standardized scenarios are still smart factories. Factories need stable output and won't pay for form factor or imagination. Over the next three years, who generates scaled revenue first — humanoid robot bodies, or smarter robotic arms? What truly sets them apart — large model capability, or force control and reliability?

Hu Xiaoping: This question is something everyone in the embodied intelligence track is watching closely. Factories are indeed the earliest automation deployment scenarios generating social value, and new technology iterations also consider rapid deployment. Early automation used industrial robotic arms — large and bulky, doing what humans couldn't, like very fast cycle times or heavy lifting. But the technical route pursued precision position accuracy, which meant it could only complete tasks in fixed scenarios and couldn't adapt to open, flexible, changing environments. Applications have limitations; industrial robot penetration remains very low, and technical breakthroughs are needed. The breakthrough direction points to humans. Factories still have many labor-intensive industries relying on human labor that industrial robotic arms can't achieve. The difference between human labor and industrial robots is in working mode. Humans are insensitive to absolute precision and position — for example, plugging in USB, you don't know the exact position but complete it every time. The perception of force and touch are etched in bones; reflecting that onto a robot is very hard. To break through automation upgrades, the force-control route is inevitable. Flexiv chose this route from the start, seeing the ability to solve complex processes on the industrial floor — both flexible assembly and complex surface treatment.

Returning to embodied intelligence. Large embodied intelligence follows large model development, hoping generalization empowers robots to solve various scenarios. But currently, the starting point is large models bringing generalization goals — hoping to do both A and B. Most factory applications require precision and stability; the two directions differ. Execution capability and manipulation capability upgrades, relative to agents, have faster priority in factory applications. We expect that for large models or embodied intelligence, factories are experimental scenarios, not current volume-scaling scenarios — what's more needed is solving robotic arm manipulation capability upgrades.

Kang Zhengzhong: I visited your company in April. Back then you mentioned adaptive robotic arms — you don't deeply feel "adaptive" without experiencing it. There was a massage robot scenario where everyone's skeleton and body type are different; many worry about inconsistent force or squeezing. After experiencing it, they use force control, vision, and AI to turn uncertainty into adaptability. You only realize "ten years of turning uncertainty into adaptability" after experiencing the product. After manipulation capability improves, combining with embodied intelligence model generalization can truly achieve the goal of robots empowering every industry.

Education Robots: Hardware Is Not a Moat

Kang Zhengzhong: Second question for WhalesBot's Mr. Jia. Your product form is hardware-centric, with robots across all age groups from 3-22, and the core may be educational programming robots. Now large models make knowledge Q&A very cheap. If an education robot just answers questions, hardware becomes a drag — an expensive shell. Compared to traditional AI learning apps, what is Pubbo's irreplaceable value? When model capability converges, is the moat hardware or content ecosystem?

Jia Xindong: People ask us every day: what's the differentiator versus other AI learning apps? Going forward, which is the moat — hardware or content?

First, WhalesBot robots focus on youth education. When young people use AI dialogue or AI software, the difference from adults is data security and content sensitivity. Adults don't care about adult content, violence, pornography, politics, or knowledge unsuitable for children — nothing is filtered. When WhalesBot decided to improve product capability in 2024, it specifically developed its own AI platform. When children use the robot, software, or content ecosystem, all Q&A goes through platform screening; things harmful to children's learning are filtered out. This is a very big differentiator.

Second, many traditional AI learning apps focus on K12 education, with underlying values still about making children study better, score higher, and master skills to show off in front of parents. This is completely different from WhalesBot's values. We want to cultivate children with AI thinking, not make them continue to compete on exams, scores, and grades. Everyone here is accomplished and has surely experienced the painful path of exams and degrees. WhalesBot's ideal is that the next generation doesn't have to go through that pain — we hope to help children find their own suitable life path through hardware and software. Not everyone has to become a scientist or programmer; we want children to find their own field.

Regarding which is the moat — hardware or content — China is extremely strong at the hardware level; no country in the world has a supply chain as deep and complete as the Pearl River Delta and Yangtze River Delta. US and European labor costs, engineer costs, and hardware production costs are far behind. Hardware cost is a very strong moat overseas. But content is actually a deeper moat. Precisely because hardware costs are already low compared to domestic competitors, if competing in international markets in the future, only doing hardware makes it hard to break out. Content is the truly long-term moat.

This content revolves around educational content. The company has been profitable for seven consecutive years since founding, with a very complete business model, surviving the pandemic and double reduction. The reason for profitability is that content is very hard-core. We had traditional programming robot content, and in 2024 started developing new AI general literacy courses to give children a good enlightenment in the AI era — interacting with hardware through natural language, programming, or other means. This is WhalesBot's true future moat.

Kang Zhengzhong: WhalesBot has raised seven funding rounds — welcome investors to connect. There's also a new business: opening 100 AI technology experience centers across China, making a Disney-like AI park for China. Friends interested are welcome to connect.

Jia Xindong: Yes, that's also our direction of effort.

Proactive Interaction: What Phones Can't Do

Kang Zhengzhong: Third question for Mr. Kong. Most agents today still "only move when you call it." You emphasize proactive interaction — requiring continuous perception or context understanding — but that also brings privacy issues and cost pressure. Must proactive intelligence rely on dedicated hardware, or can a phone plus camera suffice? Will you ultimately sell companion products, or reusable embodied intelligence brain and data capability?

Kong Weigang: Most smart hardware today, including embodied bodies, still relies on instructions or remote controls to complete tasks. Through collecting large amounts of real-scene interaction data, we build datasets to generalize proactive interaction capability.

Data value is extremely important. Phone volume is enormous, but it works in low-end or test scenarios. Truly achieving intent recognition, proactive perception, and proactive execution definitely requires new-form agent hardware or robots. Phones are designed for communication and information delivery; cameras, mic arrays, screen-off states are for that need. But the core of proactive perception — continuous perception — with the screen off, how do you continue? Intent judgment — phones can't be fixed in a scene to solve the angle problem long-term; the data source is flawed. Proactive execution — outputting instructions containing actions and language. Phones can trial-test in small scenarios, but new-form agents and robots are the primary carrier for proactive interaction — capable of continuously perceiving external changes from fixed scenarios or specific needs, forming a closed loop from perception to delivery.

Fenghuolun's strategy is to use the proactive interaction carrier as the entry point, form real data collection, and through model generalization form an overall closed loop. In the future, based on proactive interaction technology, more new-form agent hardware will emerge, helping existing agent hardware upgrade structurally and in usage scenarios.

Kang Zhengzhong: Phones may not be the best fit; dedicated hardware is easier to implement?

Kong Weigang: Correct — the carrier design starting point is completely different. Perhaps after the next major upgrade in proactive interaction technology, phone design principles will also change, aligning with next-generation interaction and technology implementation.

AI Wearables: Context Is All You Need

Kang Zhengzhong: Mr. Duan is here. Please introduce yourself.

Duan Ran: Sorry, I came from Shenzhen. Shanghai is even hotter than Shenzhen. I'm Duan Ran from Star Dome Ark, doing AI wearables, including AI glasses, pendants, rings, etc. We collaborate with upstream and downstream factories in Shenzhen, with a fairly broad product line.

Kang Zhengzhong: You mentioned doing multiple AI wearable forms — glasses, pendants, maybe also rings. Multiple forms mean resource dispersion. Beyond making money, laying out multiple wearables simultaneously — are you searching for the best carrier through different form experiments? Or do you have a bigger dream — building a multi-network agent ecosystem where devices collaborate? Also, is the AI wearable entry point ultimately determined by wearing position, or by agent capability (long-term context understanding)?

Duan Ran: Start from cost. In Shenzhen, chips uniformly use BES, JieLi, Actions, and Allwinner — unified supply this year. Cost whether multi-matrix or single-matrix — BOM is fixed; it's just a matter of shipment volume. The PCB mainly involves a few sensors: microphone sensors, NPU/VAD. The sensor supply chain doesn't change because we define an always-on context device. The supply chain doesn't change much; no new source SKUs are added — only appearance design and PCB design vary. Overall cost is not the 80% or 100% per SKU everyone imagines; it may be only about 20% in product line management. In business, the first consideration is cost control and supply chain complexity.

Is it like media — multi-matrix distribution? No. It's just that we're not sure which product users will like; under controllable cost and profit risk per product, we launch several — as long as they recoup costs, it's fine. That's the first point — very grounded.

Second, the entry point question. AI entry point isn't new; agent entry points have been discussed for two or three years, maybe one or two more. After all, no native AI hardware has broken a million units yet. Native AI hardware is defined as non-traditional large SKUs (phones, watches, earphones) — new SKUs, like AI Pin, Rabbit R1, Plaud, etc. Plaud is also a hit — a recording card plus AI features. But when will a new AI OS-native hardware appear like a computer — owned by over 70% of young people? I think maybe within three years. The specific form can't be said; if I knew, I'd be ringing the IPO bell. I can only say that wearables currently have the most potential, including glasses, or chest-worn forms like pods, pendants, pins. The core of next-generation AI hardware is context — context is all you need. Whoever can hold users' all-day context data — including offline real physical world information (who you talked to about what, where you went) and online information (what code, documents were transmitted at work). Holding a person's entire context, when base model capability is the same, will definitely perform over 40% better than other service providers — users will never leave you. Among all products, only wearable hardware can achieve this.

Kang Zhengzhong: The advantage of doing hardware in Shenzhen. Hardware changes too fast now to bet on one form long-term; hardware must also be competitive — build it first, and whichever works and has volume may be the future trend.

Brain-Computer Interface: Sell Research First, Then Wait for the Wave

Kang Zhengzhong: Last question for Mr. Li. You do brain-controlled drones — those five words are eye-catching — but from demo to product there's much to solve: wearing experience, signal stability, scenario demand. You have a new product launch in Beijing on July 20. Is this launch meant to prove technically that BCI can control, or product-wise that brain-controlled drones are worth long-term use? Is the first commercialization threshold on the technical side (algorithm accuracy) or product side (experience and trust)?

Li Yang: We do more than brain-controlled drones — we can do brain-controlled robot dogs, brain-controlled cars; drones are just a landing scenario attempt. The launch isn't to prove how good Shenwu's stuff is, but to show how to make products. We invited all BCI professors, experts, and big names. In their eyes, what truly matters is researching how to empower humans and help better understand the brain. The launch focuses on how to extract EEG more accurately. This accuracy is hard to measure by standards in daily life; only those truly doing research, brain studies, and brain cognition need EEG extraction to support further research. They can feel it — what kind of brainwave better supports research, what kind of data is high-quality. The first step is letting research experts solve laboratory rigid demand — truly extracting human brain signals, amplifying, analyzing, and decoding human brain information.

The first commercialization threshold. Collecting EEG isn't common in daily life; maybe in the future everyone wears EEG collection devices, but now it's common in laboratories and research institutions. The first threshold is how to make collection more accurate and convenient. Most companies started from laboratories; truly studying EEG, every collection was complex and painful. Our product makes EEG collection simpler, more efficient, and more accurate.

Duan Ran: I'm quite interested in EEG. There are many EEG sleep devices and EEG pillows now under various banners. Mr. Li is formally trained — I'd like a layperson's explanation: do the non-invasive brain rings and EEG pillows currently on the market truly collect accurate data? Can they help users in motion or sleep with suggestions? Recently a brain ring company got Shenzhen funding — what do you think?

Li Yang: Shenwu also does non-invasive. Brain ring products haven't launched yet; we've researched for nearly a year. From an EEG collection perspective, how accurate and high-density can non-invasive data be for human use? Head-mounted headbands are mostly dual-channel, collecting from the prefrontal cortex — considered the thinking, active signal area. They can collect brain activity; technically, our product can definitely achieve this — our accuracy is very high. For example, right now in this conversation, brain activity must be very high. When I just drifted off while listening to other guests, attention dropped — that's also easily collected. Feedback into sleep scenarios: lying in bed at night, whether you're ruminating or entering the sleep process, it can be accurately collected. After accurate data collection, with feedback loop methods, influencing sleep posture, lighting, sound, surrounding environment — once product safety and standards are in place and other EEG signals are added, theoretically it can be achieved.

Beyond the prefrontal cortex, how do brain-controlled drones work? They collect signals from the visual cortex in the back of the head, based on steady-state visual evoked potentials — collecting the frequencies the human eye sees. External fixed frequencies are emitted, and the collection device recognizes what you're looking at — that's how it works. The brain-controlled typing and brain-controlled smart terminals discussed in lab news before are all possible.

BCI is like a base — a new human-machine interaction mode that can empower smart wearables and AI hardware. Smart hardware like personal friends can more directly understand brain emotions, focus, and activity. Once the industry later achieves EEG large models, AI devices can read what their partner is thinking — and it might also empower WhalesBot's education industry.

Jia Xindong: Brain-controlled drones sound sexy. We also want to read and identify students' focus issues — to solve them. Focus can also be accurately collected.

Kang Zhengzhong: It looks like you're currently going B2B first, selling to professionals or research institutions. The earliest commercial closed loop — is it selling hardware, data services, or training solutions?

Li Yang: Hardware first. Without a hardware interface, EEG data can't be collected; the hardware must be precise, accurate, and high-quality. The first step is universities and research institutions — they need to recognize the equipment, that the collected data is usable and high-quality. Now Tsinghua, Peking University, and major universities across China are all building BCI disciplines, using our equipment for EEG experiments. Only after recognition is the data high-quality and usable. We're also accumulating data. Data services will definitely be the inevitable path for the BCI or EEG large model industry; whoever first occupies data quality and builds a data platform truly occupies the upstream.

Commercialization: From G to B to C Path

Kang Zhengzhong: Second perspective — commercialization. Mr. Jia, education robot users have two decision-makers: whether the child wants to use it, and whether the parent will pay — these don't equate. Over seven years of consecutive profitability, did you initially target families, schools, or competitions? Is there data proving buyers get long-term value, not a one-time novelty experience?

Jia Xindong: WhalesBot's B2B and G2B models are already mature. The competition side is very closed-loop — we self-host the ENJOY AI competition, a global youth AI robotics innovation competition, running for over seven consecutive years, with over 150,000 teenagers participating annually from around the world; regional winners come to China for finals. Previous years were in Shanghai, in the internet town; this year in Hengdian. Competition profitability has been a very important part.

Second, educational robots — over the past few years and after this year's 15th Five-Year Plan, the country attaches great importance. A previously important piece was cooperation with government, schools, and institutions. The government pushes AI education; many teachers and principals don't know how to teach, have no teaching tools, hardware, or curriculum content. WhalesBot used to provide solutions. Currently partnering with over 30,000 schools domestically, over 10,000 institutions are distributors and beneficiaries. The second profit center was to G.

On the family side, whether a child likes it and parents pay are indeed two different things. This is WhalesBot's next-step plan. Competition and government channels and business models are mature; the next step is entering families and the C-end market — this is one of my main directions.

Winning the child or the parent? We want to achieve educational goals for children's future, letting parents know what their child is suited for. Second, the AI era education industry is changing; we hope to solve parental anxiety: what should children learn, how to teach children to use AI, what to avoid in the AI process. These form long-term usage value that wins parents.

For example, Imoo watch — we've been benchmarking and learning from it. It's the brand that found a breakthrough between parents and children. On the parent side, safety positioning, knowing where the child is in real time, solving safety anxiety. On the child side, watch design, social features, entertainment — children love using it. Future WhalesBot C-end products and existing robot products will continue following Imoo's path, hoping to become the next generation of Imoo.

Kang Zhengzhong: From G to B to C, each approach is different.

Jia Xindong: The future vision is not just selling hardware or content, but more importantly, data. In the AI era, more people connect to the data network. WhalesBot has deeply cultivated youth technology for seven years; the data we have and will continue to accumulate will be very valuable — this is also the future commercialization direction.

Industrial Client ROI and Payback Period

Kang Zhengzhong: Third question for Mr. Hu. Industrial clients don't pay for generality or form factor itself; they care more about yield and payback period. You mentioned adaptability getting more general, but client delivery getting more customized. How do you turn "flexibility" into quantifiable ROI? If you can only use one metric to drive a purchase — yield, cycle time, or payback period?

Hu Xiaoping: From a manufacturing enterprise's selection perspective, yield and reliability are prerequisites, not evaluation criteria — if you don't meet them, they might not even consider you. What ultimately drives the transaction is definitely ROI and cost recovery period — every industry is the same. Many manufacturing businesses have very thin margins; payback period is a key consideration.

Industrial manufacturing pursues reliability and stability, but turning uncertainty into adaptability or generality is more about bringing down fixed-scene or deployment costs. In the past, to adapt to controlled environments, you had to work with many peripheral non-standard devices, greatly increasing operational and maintenance costs. Once robots can autonomously adapt to environmental changes, non-standard devices can be greatly reduced, and the overall solution improves qualitatively, encouraging more industries to use our arms. After deployment, not just automotive and smartphone industries but even food processing, medical, laboratory automation, and other industries previously done by humans or unimaginable for automation have seen breakthroughs. This is a typical scenario of technology-driven industrial upgrading.

Kang Zhengzhong: Follow-up question. Future globalization — non-standard gradually becomes standardized processes, replicated to Europe or North America. Do you prioritize replicating standard processes, or entering overseas factories with leading overseas clients first?

Hu Xiaoping: We do both. First, build flagship cases with leading enterprises — the driving effect is large enough to influence the entire related industry chain's understanding of the solution. On this basis, build standardized platforms or solutions, improving product ease of use and peripheral support. This enables larger-scale, faster application replication. For example, after landing with a Shanghai or domestic Fortune 500 company, they quickly推广 the solution to North American, European, and Southeast Asian factories. It has both benchmark effect and accelerates the growth curve — mutually reinforcing.

Giants Everywhere: How Do Startups Survive?

Kang Zhengzhong: Mr. Duan, many giants have joined AI wearables — Google, Huawei, Alibaba. Startups can't win head-on on specs. You mentioned starting with narrow scenarios, but narrow scenarios don't have a strong volume-scaling narrative. For the first overseas market, do you enter vertically first, or directly build a micro-innovation new-category consumer market?

Duan Ran: Giants everywhere: Xiaomi, Huawei, OPPO, Vivo, Honor, Meta, Apple, and a string of big companies. How do you find a way to survive in a red ocean? First, survive. Even Realities is worth referencing — they make AR optical glasses and raised about 1 billion RMB two months ago, selling well in Europe. First, find where giants can't reach. Back then everyone was doing full-color or monochrome green waveguides plus mics and speakers — big and comprehensive. They targeted European business professionals, removed the speaker, kept only the mic and display — a lightweight approach entering a relatively high-net-worth segment, and succeeded. So first, find your own positioning. Second, even a narrow track, if you go global from day one, can find a large enough base in the global market. There are no narrow tracks, only your narrow field of vision.

Second, you can also hold onto a giant's coattails. New product launches at Huawei's HC Full Connect Conference in September. Whether it's Huawei's agent secondary development platform or its HarmonyOS ecosystem, it brings a huge user base. Fundamentally, holding onto a giant's coattails is also part of the strategy. If you don't have to change the world like Steve Jobs, being like Xiaomi is fine. First leverage the existing supply chain, cloud platform agent compute, and base models to build your own product. As long as you make it faster, cheaper, better, and solve pain points for users, you can still build differentiation. Even Xiaomi didn't make the first phone or the first car, not even the first five — but it's still a strong top-three contender in the current market.

Kang Zhengzhong: Understood. The market is big enough; elephants don't crush ants — it's not zero-sum. Make differentiation, partner with big companies, grow the pie.

Proactive Interaction Hardware Overseas

Kang Zhengzhong: Last question for Mr. Kong. Proactive interaction dedicated hardware entering overseas homes — language is the first barrier, plus sensitive words, minors, family data, companion boundaries; the trust threshold keeps rising. Chinese scenarios may not replicate overseas. Validate the market with consumer products first, or directly export embodied intelligence brain capability to overseas giants? For the first sample market, prove shipment volume first, or prove long-term retention value?

Kong Weigang: One is hardware, one is the brain. The brain's generalization core comes from data. Only when underlying data is sufficient, with enough real-world scenario data, does generalization happen — like our proactive interaction brain built on our own dataset, achieving generalization with small data volume. The data source — we want to establish a commercial closed loop through bulk hardware delivery; that's our route. Through our self-developed device "Xiao Jing" — a new-form agent hardware for home learning and growth, launched on the 15th of last month — it's now供不应求. Why choose this path? Solving real user needs — only real demand drives real commercial scenario deployment, which solves the data problem and forms high-frequency use. This closed loop lets the proactive interaction brain form data feedback. Overseas markets proceed with the same theory.

The difference between domestic and overseas is localization policy and personal information protection laws — completely different, even stricter. Bulk hardware delivery with data feedback through hardware works very well. Conversely, exporting the proactive interaction brain first and combining with others' hardware creates gaps in hardware adaptability and scenario integration, leading to poor user experience and unable to drive high-frequency usage scenario data. So both overseas and domestic, we choose real-scenario agent hardware to solve core needs and form a commercial closed loop.

Kang Zhengzhong: Summarizing, your data collected through the real physical world is also a unique advantage compared to internet corpora?

Kong Weigang: Correct. The proactive interaction brain differs from large language models and multimodal models — through processing causal data from the real world, the brain understands why it does something, the process, and the final result, forming a complete closed loop through this underlying data.

Originally published by Unique Research on Unique Research Substack on August 5, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research