arXiv:2608.23189cs.CV2026-08

EchoWM实现可进入的多模态世界模型,支持连续导航下的音视频同步生成。

EchoWM: Open and Enterable Omnimodal World Models

论文配图:EchoWM: Open and Enterable Omnimodal World Models
图 1 · 摘自论文原文
  • 基于相机意图统一建模第一/第三人称视角,实现6自由度连续运动控制
  • 联合生成720p视频、环境音、音乐和语音,长时序生成保持同步性
  • 适用于多种场景的交互式生成,适合虚拟世界构建与沉浸式应用

我们提出EchoWM,一种可进入的多模态世界模型,能够在连续导航过程中协同生成720p视频、环境音、音乐和语音。交互以相机意图为核心:第一人称场景中定义观察者运动,第三人称场景中通过数据学习相机-角色动态关系,无需特定视角控制器。离散指令与连续位姿映射到共享度量尺度的相对6-DoF轨迹,数据集级校准确保异构数据间运动幅度一致。为联合学习音频-视觉生成与轨迹控制,我们构建互补数据引擎,并采用渐进式训练结合自回归后训练,实现长时序生成。大量评估表明,该模型在公开世界模型基准上表现优异,具备强轨迹跟随能力与高视觉质量,支持跨主题的第一/第三人称交互,并在长时生成中保持环境音与语音的同步性。

原文摘要 · Abstract (English)

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.

多模态生成世界模型音视频同步可进入内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。