arXiv:2604.09473cs.CV2026-04被引 1

构建可自由移动的沉浸式立体视频,实现音画同步的6自由度交互体验。

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

  • 基于高斯时空表示与多视角音视频数据,实现动态光场重建。
  • 首次从多视角音视频中恢复声场,支持真实空间感音频反馈。
  • 提供5K/60fps高清数据集,适用于复杂场景的虚实融合应用。

完全沉浸式的虚拟与增强现实体验需要紧密集成6自由度视觉与听觉交互。尽管可通过计算机生成内容实现,但直接从真实世界视频构建此类体验仍鲜有研究。本文提出沉浸式立体视频(Immersive Volumetric Videos, IVV),一种支持大范围6-DoF交互、音画反馈及高分辨率高帧率动态内容的新立体媒体格式。为支持IVV构建,我们推出基于空间导向采集理念的多视角多模态数据集ImViD,自研采集装置可在运动中同步获取多视角音视频,高效捕获包含丰富前景-背景交互和复杂动态的室内外场景。数据集提供5K分辨率、60帧/秒、持续1-5分钟的视频,其空间、时间与多模态覆盖优于现有基准。基于此,我们开发了基于高斯的动态光场重建框架,结合流引导稀疏初始化、联合相机时间校准与多目标时空监督,实现复杂运动的鲁棒精准建模。此外,我们首次提出从多视角音视频数据中重建声场的方法。上述组件构成统一的沉浸式立体视频生成流水线。大量基准测试与沉浸式VR实验表明,该方案可生成高质量、时序稳定的音画立体内容,支持大范围6-DoF交互。本工作既定义了沉浸式立体视频的基础范式,也提供了可行的构建方法。

原文摘要 · Abstract (English)

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constructing them directly from real-world captured videos remains largely unexplored. We introduce Immersive Volumetric Videos, a new volumetric media format designed to provide large 6-DoF interaction spaces, audiovisual feedback, and high-resolution, high-frame-rate dynamic content. To support IVV construction, we present ImViD, a multi-view, multi-modal dataset built upon a space-oriented capture philosophy. Our custom capture rig enables synchronized multi-view video-audio acquisition during motion, facilitating efficient capture of complex indoor and outdoor scenes with rich foreground--background interactions and challenging dynamics. The dataset provides 5K-resolution videos at 60 FPS with durations of 1-5 minutes, offering richer spatial, temporal, and multimodal coverage than existing benchmarks. Leveraging this dataset, we develop a dynamic light field reconstruction framework built upon a Gaussian-based spatio-temporal representation, incorporating flow-guided sparse initialization, joint camera temporal calibration, and multi-term spatio-temporal supervision for robust and accurate modeling of complex motion. We further propose, to our knowledge, the first method for sound field reconstruction from such multi-view audiovisual data. Together, these components form a unified pipeline for immersive volumetric video production. Extensive benchmarks and immersive VR experiments demonstrate that our pipeline generates high-quality, temporally stable audiovisual volumetric content with large 6-DoF interaction spaces. This work provides both a foundational definition and a practical construction methodology for immersive volumetric videos.

立体视频6-DoF音画同步多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。