arXiv:2411.11199cs.CVeess.IV2024-11被引 3

开源多视角人体体视频数据集,助力三维视频压缩研究

BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression

  • 构建18组多视角RGB-D数据与对应网格模型
  • 神经编码方法相比传统方法平均提升38%码率性能
  • 适合三维重建、压缩与质量评估方向研究者使用

沉浸式技术和3D重建的进步使得高精度数字复制品得以实现,但随之产生的海量3D数据亟需高效压缩以应对存储与传输的带宽限制。然而,现有高质量体视频数据集稀缺,采集成本高且资源消耗大。为此,我们提出开源多视角人体体视频数据集BVI-CR,包含18组多视角RGB-D捕获及其对应的纹理多边形网格,涵盖多样人体动作。每段视频含10个视角,分辨率为1080p,时长10-15秒,帧率30FPS。基于MPEG MIV通用测试条件,我们对三种传统及基于神经坐标系的多视角视频压缩方法进行了基准测试,结果显示神经表示方法在体视频压缩中潜力显著,相较于传统方法在PSNR上平均提升达38%。该数据集为体重建、压缩与质量评估提供了统一平台。数据将公开共享于https://github.com/fan-aaron-zhang/bvi-cr。

原文摘要 · Abstract (English)

The advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38\% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at \url{https://github.com/fan-aaron-zhang/bvi-cr}.

体视频数据集压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。