arXiv:2512.00883cs.MMcs.CV2025-12

构建音视频世界模型,让智能体能同时预测视觉与听觉变化。

Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics

  • 将音视频同步观测建模为部分可观测马尔可夫决策过程。
  • 在76个室内场景的30小时数据上实现高保真视听预测。
  • 适合做多模态感知、具身导航的研究者参考。

世界模型通过模拟环境动态,使具身智能体能够规划和推理未来状态。尽管真实感知天然多模态,现有方法主要聚焦视觉观测,忽视了空间和时间上的声学线索。本文提出统一的音视频世界模型(AVWM),将动作控制下的多模态环境模拟形式化为带有同步音视频观测的部分可观测马尔可夫决策过程。作为基础基准,我们构建了AVW-4k数据集,包含76个室内环境的30小时带动作标注的双耳音视频轨迹。为捕捉物理驱动的多感官动态,我们提出AV-CDiT(音视频条件扩散变换器),采用新颖的模态专家架构平衡视觉与听觉学习,并通过三阶段训练策略优化。大量实验表明,AV-CDiT在视觉与听觉两个模态上均实现高保真预测。此外,我们在具身导航任务中验证其实用性,结果表明AVWM显著提升预训练智能体在连续音视频导航中的表现。

原文摘要 · Abstract (English)

World models simulate environmental dynamics to enable embodied agents to plan and reason about future states. While real-world perception is inherently multimodal, existing approaches focus primarily on visual observations, leaving crucial spatial and temporal acoustic cues underexplored. In this work, we present a unified formulation of Audio-Visual World Models (AVWM), casting multimodal environment simulation under action control as a partially observable Markov decision process with synchronized audio-visual observations. As a foundational benchmark, we construct AVW-4k, comprising 30 hours of action-annotated binaural audio-visual trajectories across 76 indoor environments. To capture these physically grounded multisensory dynamics, we propose AV-CDiT (Audio-Visual Conditional Diffusion Transformer), featuring a novel modality expert architecture that balances visual and auditory learning, optimized via a three-stage training strategy. Extensive experiments demonstrate that AV-CDiT achieves high-fidelity prediction across both visual and auditory modalities. Furthermore, we validate its practical utility in embodied navigation, showing that AVWM significantly enhances a pretrained agent in continuous audio-visual navigation tasks.

多模态世界模型具身智能音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。