构建真实遮挡下的3D人体姿态与形状估计数据集,填补现实场景评估空白。
VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real Occlusions
- 基于计算机图形学生成包含复杂遮挡的真实视频数据
- 在多个公开数据集上显著提升模型性能,验证方法有效性
- 适用于研究遮挡下人体检测与姿态估计的学者与工程师
人体姿态与形状(HPS)估计方法虽在自然图像和视频中表现优异,但在复杂姿态或严重遮挡场景中仍面临挑战。现有数据集多采用随机补丁或剪贴画式遮挡,难以反映真实世界问题。为此,我们提出新型视频基准数据集VOccl3D,包含3D人体姿态与形状标注,通过先进渲染技术模拟多样真实遮挡、服装纹理与动作。我们在该数据集上微调最新方法CLIFF与BEDLAM-CLIFF,实现多个公开数据集及本数据集测试集上的显著提升,并利用其优化物体检测器YOLO11,构建端到端鲁棒的遮挡下HPS系统。该数据集为评估抗遮挡方法提供更真实的基准,推动相关研究发展。
原文摘要 · Abstract (English)
Human pose and shape (HPS) estimation methods have been extensively studied, with many demonstrating high zero-shot performance on in-the-wild images and videos. However, these methods often struggle in challenging scenarios involving complex human poses or significant occlusions. Although some studies address 3D human pose estimation under occlusion, they typically evaluate performance on datasets that lack realistic or substantial occlusions, e.g., most existing datasets introduce occlusions with random patches over the human or clipart-style overlays, which may not reflect real-world challenges. To bridge this gap in realistic occlusion datasets, we introduce a novel benchmark dataset, VOccl3D, a Video-based human Occlusion dataset with 3D body pose and shape annotations. Inspired by works such as AGORA and BEDLAM, we constructed this dataset using advanced computer graphics rendering techniques, incorporating diverse real-world occlusion scenarios, clothing textures, and human motions. Additionally, we fine-tuned recent HPS methods, CLIFF and BEDLAM-CLIFF, on our dataset, demonstrating significant qualitative and quantitative improvements across multiple public datasets, as well as on the test split of our dataset, while comparing its performance with other state-of-the-art methods. Furthermore, we leveraged our dataset to enhance human detection performance under occlusion by fine-tuning an existing object detector, YOLO11, thus leading to a robust end-to-end HPS estimation system under occlusions. Overall, this dataset serves as a valuable resource for future research aimed at benchmarking methods designed to handle occlusions, offering a more realistic alternative to existing occlusion datasets. See the Project page for code and dataset:https://yashgarg98.github.io/VOccl3D-dataset/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。