When AI can compose poetry, paint pictures, and generate hyper-realistic videos, the next frontier is making it truly capable of “getting its hands dirty”—understanding the physical laws of the real world and interacting with its environment in real time. At the recently concluded 2026 World Artificial Intelligence Conference (WAIC) in Shanghai, this proposition became the industry-wide consensus, with the “world model” thrust into the spotlight and viewed by multiple Chinese firms as the strategic high ground for the next phase of AI competition.
“2026 is the ‘Year One of the World Model,'” declared Fang Han, Chairman and CEO of Kunlun Tech, during the conference. He argued that AI’s core challenge has officially shifted from the static content generation of the past two years—text, images, video—toward “understanding” and “interaction,” meaning enabling AI to genuinely grasp the operating principles of the physical world and engage in closed-loop interaction with real environments. The world model, he said, is the critical technological foundation for achieving this leap.
Judging by the exhibitor lineup at this year’s WAIC, Chinese enterprises are markedly accelerating R&D in the world model domain, with a diverse array of approaches on display. Kunlun Tech unveiled its Matrix-Game 3.5 world model, which focuses on interactive technology and innovatively implements patch-level memory injection—splitting images or video frames into tiny localized blocks with spatial coordinates—to ensure long-term consistent spatial memory. The model has built an automated data production pipeline covering over 1,200 game scenarios, and the company announced that its core architecture will be open-sourced.
Ant Lingbo introduced LingBot-VA 2.0, an embodied native world action model that takes a different technical path: rather than grafting digital world model capabilities onto a robot’s “brain,” it is natively designed from the ground up based on the original requirements of interacting with the environment, including dynamic modeling, causal prediction, and real-time execution. Elsewhere, Jijia Vision showcased for the first time its full suite of “universal world model” products spanning from generation to action and from the robot body to the scene. Zhiyuan exhibited its GE-2 world model, which enables robots to continuously learn from environmental changes, providing a closed-loop environment for iterative improvement. Daxiao Robotics released its Kaiwu World Model 3.1, which integrates generative intelligence, physical intelligence, and cognitive intelligence.
Notably, 3D generation companies are also accelerating their entry into the race to solve the problem of rendering and simulating the world. Yingmo Technology’s Hyper3D WorldGen made its debut, marking a shift in 3D generation from individual objects to the “scene level,” achieving end-to-end conversion from real imagery to trainable simulation scenes, with more 3D objects becoming directly usable for robotic simulation training.
Although the concept of the “Year One of the World Model” has been floated, the industry is far from reaching a unified technical consensus. The “world model” umbrella is broad, and companies differ significantly in their technical approaches and positions along the industry chain, with each beginning to bet on different paths.
Fang Han’s assessment is that development will initially proceed through Scaling Laws—where model capability improves as data volume, computing power, and parameter count increase—and technologies will eventually converge. “All paths will ultimately converge, just as autonomous driving went through a contest of multiple technical routes before eventually converging on end-to-end models, but convergence takes time.” He explicitly stated that at least 10 million hours of training data are needed before model capability can potentially reach a qualitative inflection point, and in some cases, as much as 100 million hours may be required.
In exploring ways to reduce data costs, Fang Han is bullish on heterogeneous data as the primary source for pre-training world models or embodied models. Traditional motion capture collection is extremely expensive: a single professional capture equipment set costs over 100,000 yuan (approximately $14,810), with per-hour collection costs reaching around 1,000 yuan, meaning assembling a 10-million-hour sample set would require investments on the order of several billion yuan—a burden most enterprises cannot bear. By contrast, using heterogeneous data such as everyday operation videos captured by smartphones and ordinary cameras can bring per-hour collection costs down to single-digit yuan, with diversity far exceeding that of professional motion capture footage. He revealed that Kunlun Tech’s Riemann 1.0 embodied model, trained on massive heterogeneous data, delivered an 8-percentage-point performance improvement on complex household tasks such as folding clothes.
Chen Yilun, founder and CEO of Tashizhihang, highlighted the current challenges facing world models from another dimension. He stressed that the biggest challenge is not generating a world that “looks real,” but truly understanding the physical laws governing the real world. “Relying solely on video data cannot capture the complete laws of the world. Vision can provide geometric and semantic information, but cannot directly obtain critical physical information such as force and contact.” For robots that need to manipulate and interact, multimodal data including tactile and force sensing is indispensable; what is fundamentally needed is real-world “work data.”
On the competition surrounding data production capacity, Fang Han believes Chinese firms hold a natural advantage. “China is the manufacturing hub for all data collection equipment hardware, and has advantages in data collection tools including domestic services and other application scenarios. I believe the largest data production capacity lies in China.” He predicted that by the end of 2026, some players could reach the million-hour data scale; by the end of next year, some may hit the 10-million-hour mark; and reaching the 100-million-hour level will require at least three to five years.
In terms of application prospects for world models, the gaming industry is widely seen as the first sector to be transformed. Fang Han noted that as world models evolve, open-world games will no longer depend on hundreds of people manually building content over several years; virtual worlds will be able to grow and evolve continuously in real time based on player actions. Within the next three to five years, world models will become infrastructure for the gaming industry, fundamentally reshaping content production methods. In the more imaginative realm of embodied intelligence, the industry is still evolving rapidly. There remains debate over whether the endgame for physical AI is an end-to-end VLA (Vision-Language-Action model), a world model, a world action model, or a more complex agent system, with many players simultaneously betting on multiple paths.
Regarding when the “ChatGPT moment” for embodied intelligence will arrive, predictions from multiple industry figures at WAIC ranged from two to five years. The consensus, however, is that the emergence of intelligence in physical AI will not come from a single model, a single data source, or a single technical path, but will depend on a continuously evolving flywheel composed of data scale, model architecture, simulation evaluation, real-machine feedback, deployment scenarios, and hardware systems.
While the world model was the technical star of this year’s WAIC, the commercialization of multimodal AIGC also delivered scaled results. According to a report by Time Weekly, data shared by Yang Haitao, Senior Vice President of iQiyi, showed that the platform now launches over 3,000 AI short dramas and more than 1,000 AI comic content pieces daily. AI tools have boosted the efficiency of pre-production storyboarding and post-production visual effects by over 50%, and can save three weeks of production time for a single online film.
Fang Han assessed that the growth pace of large text models is gradually slowing, but the three major tracks of video, music, and gaming remain in a clear performance dividend phase, with their commercialization ceilings far from being reached. Kunlun Tech’s newly upgraded music large model, Mureka v9.5, carries the label of being “the least AI-sounding,” with instruction precision improved by over 10%, and has connected the pipeline for one-click soundtracking of images, text, and video, supporting collaborative music video creation across the entire workflow.
Addressing the impact of AI technology on entertainment industry professionals, Fang Han drew an analogy to “the panic of carriage drivers when the automobile was invented,” arguing that those who first master new technologies will be the first to reap the dividends of the era. Huang Xiaoming, Vice Chairman of the China Film Association, and Chinese mainland actress Wang Luodan, who shared their creative experiences on site, expressed similar views, describing AI as a highly efficient tool while maintaining that the core of storytelling and aesthetic judgment will always be dominated by humans. As the barriers to creation continue to fall, unleashing the creativity of more ordinary people, the entire AIGC industry will embrace broader growth opportunities.
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.
The news and data on this website are for reference only and do not constitute investment advice or an offer to buy or sell. Information is sourced from exchanges and public sources, and may be delayed, interrupted, or updated. While we strive for accuracy, we do not guarantee timeliness, correctness, or completeness. Content may include external links for which we are not responsible. By using this site, you agree that we and our partners are not liable for any losses. Investment carries full responsibility; please carefully assess risks and consult professionals. If there are errors in the content, please contact us for correction.
AI Search


