arXiv:2409.12980cs.CV2024-09被引 1

构建30视角的多人物-物体交互数据集,支持高质量三维重建研究。

A New People-Object Interaction Dataset and NVS Benchmarks

  • 用30台Kinect Azure采集多视角RGB-D视频,每帧4K分辨率
  • 包含38组序列,时长1~19秒,附带相机参数与人体模型
  • 提供完整标注,适合研究人-物交互与新视图合成任务

近期,人-物交互场景下的新视图合成(NVS)受到越来越多关注。现有交互数据集多为静态、视角有限的RGB图像或视频,主要包含单人与物体的交互,且存在光照复杂、同步性差、分辨率低等问题,限制了高质量人-物交互研究的发展。本文提出一个新的多人-物交互数据集,包含38组30视角的多/单人RGB-D视频序列,每段视频由30台统一布置的Kinect Azure设备捕获,分辨率为4K,帧率25 FPS,持续时间1~19秒。数据集还包含相机参数、前景掩码、SMPL人体模型、部分点云和网格文件。同时,我们在该数据集上评估了若干SOTA NVS模型,建立了新的NVS基准。希望本工作能推动人-物交互领域的进一步研究。

原文摘要 · Abstract (English)

Recently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing interactions between a single person and objects. Moreover, these datasets exhibit complexities in lighting environments, poor synchronization, and low resolution, hindering high-quality human-object interaction studies. In this paper, we introduce a new people-object interaction dataset that comprises 38 series of 30-view multi-person or single-person RGB-D video sequences, accompanied by camera parameters, foreground masks, SMPL models, some point clouds, and mesh files. Video sequences are captured by 30 Kinect Azures, uniformly surrounding the scene, each in 4K resolution 25 FPS, and lasting for 1$\sim$19 seconds. Meanwhile, we evaluate some SOTA NVS models on our dataset to establish the NVS benchmarks. We hope our work can inspire further research in humanobject interaction.

人-物交互新视图合成多视角数据集RGB-D

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。