构建可实时交互的4D动态世界模型,实现持久记忆与长期一致性生成。
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
- 通过生成-重建-引导闭环,统一视频生成与场景重建
- 支持长时序生成且延迟低,实现在实际算力下的实时合成
- 适合需要持续感知与交互的智能体、虚拟环境构建场景
世界模型旨在赋予人工智能系统以连贯、时间一致的方式表征、生成和交互动态环境的能力。尽管近期视频生成模型在视觉质量上表现优异,但在实时交互、长时序一致性和动态场景持久记忆方面仍受限,难以演变为实用的世界模型。本文提出TeleWorld,一个实时多模态4D世界建模框架,将视频生成、动态场景重建与长期世界记忆整合于闭环系统中。TeleWorld引入新颖的生成-重建-引导范式:生成的视频流被持续重构为动态4D时空表示,进而指导后续生成以维持空间、时间与物理一致性。为支持长时序生成并保持低延迟,采用基于自回归扩散的视频模型,并结合宏观-微观规划(MMPL)——一种从帧级到段级降低误差累积的分层规划方法,辅以高效的分布匹配蒸馏(DMD),实现在实际计算预算下的实时合成。本方法在统一4D框架内实现了动态物体建模与静态场景表示的无缝融合,推动世界模型向可交互、具记忆、计算可行的方向迈进。大量实验表明,TeleWorld在静态与动态世界理解、长期一致性及实时生成效率方面均表现优异,是迈向多模态生成与具身智能中交互式、记忆型世界模型的重要一步。
原文摘要 · Abstract (English)
World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual quality, they remain limited in real-time interaction, long-horizon consistency, and persistent memory of dynamic scenes, hindering their evolution into practical world models. In this report, we present TeleWorld, a real-time multimodal 4D world modeling framework that unifies video generation, dynamic scene reconstruction, and long-term world memory within a closed-loop system. TeleWorld introduces a novel generation-reconstruction-guidance paradigm, where generated video streams are continuously reconstructed into a dynamic 4D spatio-temporal representation, which in turn guides subsequent generation to maintain spatial, temporal, and physical consistency. To support long-horizon generation with low latency, we employ an autoregressive diffusion-based video model enhanced with Macro-from-Micro Planning (MMPL)--a hierarchical planning method that reduces error accumulation from frame-level to segment-level-alongside efficient Distribution Matching Distillation (DMD), enabling real-time synthesis under practical computational budgets. Our approach achieves seamless integration of dynamic object modeling and static scene representation within a unified 4D framework, advancing world models toward practical, interactive, and computationally accessible systems. Extensive experiments demonstrate that TeleWorld achieves strong performance in both static and dynamic world understanding, long-term consistency, and real-time generation efficiency, positioning it as a practical step toward interactive, memory-enabled world models for multimodal generation and embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。