构建首个可移动多视角音视频数据集,推动高沉浸感虚拟现实体验
ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
- 提出可移动多视角音视频采集系统,支持动态场景完整捕获
- 数据集含5K分辨率60帧/秒视频,时长1-5分钟,覆盖室内外复杂场景
- 提供基准方案,助力6-DoF多模态沉浸式虚拟现实研究
用户参与度因融合视觉与听觉的全沉浸多模态体验而显著提升。因此,下一代VR/AR技术的关键在于实现完全场景捕捉、大6-DoF交互空间、多模态反馈及高分辨率、高帧率内容的沉浸式体素视频。为推动该领域发展,我们提出ImViD,一个面向空间感知的多视图、多模态数据集,涵盖多种室内外场景。我们的采集装置支持移动状态下的多视角音视频同步捕获,解决了现有数据集的局限,极大提升了数据采集的完整性、灵活性与效率。捕获的多视角视频(配同步音频)分辨率达5K,帧率为60FPS,持续1-5分钟,包含丰富的前景背景元素与复杂动态。我们基于该数据集对现有方法进行基准测试,建立从多视角音视频输入构建沉浸式体素视频的基线流程,用于6-DoF多模态沉浸式VR体验。基准结果与重建交互表现证明了数据集与基线方法的有效性,我们相信这将激发未来沉浸式体素视频制作的研究。
原文摘要 · Abstract (English)
User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interaction space, multi-modal feedback, and high resolution & frame-rate contents. To stimulate the reconstruction of immersive volumetric videos, we introduce ImViD, a multi-view, multi-modal dataset featuring complete space-oriented data capture and various indoor/outdoor scenarios. Our capture rig supports multi-view video-audio capture while on the move, a capability absent in existing datasets, significantly enhancing the completeness, flexibility, and efficiency of data capture. The captured multi-view videos (with synchronized audios) are in 5K resolution at 60FPS, lasting from 1-5 minutes, and include rich foreground-background elements, and complex dynamics. We benchmark existing methods using our dataset and establish a base pipeline for constructing immersive volumetric videos from multi-view audiovisual inputs for 6-DoF multi-modal immersive VR experiences. The benchmark and the reconstruction and interaction results demonstrate the effectiveness of our dataset and baseline method, which we believe will stimulate future research on immersive volumetric video production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。