arXiv:2506.00558cs.CV2025-06被引 1

首个兼顾人体特征与动态视觉效果的体素视频数据集,助力重建与压缩算法研究。

ViVo: A Dataset for Volumetric Video Reconstruction and Compression

  • 构建真实生产流程下的多视角RGB/深度视频对,含同步音频与3D点云。
  • 包含14组多视图数据,每帧30帧率,带逐帧标定与前景掩码。
  • 适用于体素视频重建、压缩算法评测,尤其适合关注复杂材质的研究者。

随着神经体素视频重建与压缩研究的发展,亟需多样且真实的基准数据集以支持模型开发与验证。然而现有体素视频数据集在语义与低层特征上缺乏多样性,难以反映真实生产流程中的复杂场景。为此,我们提出首个扩展多样性定义的体素视频数据集ViVo,涵盖人体特征(如皮肤、头发)与动态视觉现象(如透明、反光、液体等)。每个视频序列包含14组同步多视角RGB与深度视频对,采样频率为30FPS,附带逐帧校准信息、音频数据、2D前景掩码及3D点云。为验证其有效性,我们基准测试了三种先进3D重建方法和两种体素视频压缩算法,结果表明该数据集具有挑战性,暴露了现有数据集在重建与压缩任务中的局限性,凸显了开发更高效算法的必要性。数据集与评测结果详见https://vivo-bvicr.github.io/

原文摘要 · Abstract (English)

As research on neural volumetric video reconstruction and compression flourishes, there is a need for diverse and realistic datasets, which can be used to develop and validate reconstruction and compression models. However, existing volumetric video datasets lack diverse content in terms of both semantic and low-level features that are commonly present in real-world production pipelines. In this context, we propose a new dataset, ViVo, for VolumetrIc VideO reconstruction and compression. The dataset is faithful to real-world volumetric video production and is the first dataset to extend the definition of diversity to include both human-centric characteristics (skin, hair, etc.) and dynamic visual phenomena (transparent, reflective, liquid, etc.). Each video sequence in this database contains raw data including fourteen multi-view RGB and depth video pairs, synchronized at 30FPS with per-frame calibration and audio data, and their associated 2-D foreground masks and 3-D point clouds. To demonstrate the use of this database, we have benchmarked three state-of-the-art (SotA) 3-D reconstruction methods and two volumetric video compression algorithms. The obtained results evidence the challenging nature of the proposed dataset and the limitations of existing datasets for both volumetric video reconstruction and compression tasks, highlighting the need to develop more effective algorithms for these applications. The database and the associated results are available at https://vivo-bvicr.github.io/

体素视频数据集三维重建压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。