arXiv:2410.18723cs.CVcs.HC2024-10被引 4

提出首个可泛化到新数据集与新关键点的多视角多人姿态估计方法。

VoxelKeypointFusion: Generalizable Multi-View Multi-Person Pose Estimation

  • 基于体素与关键点融合,提升多视角多人姿态估计泛化能力。
  • 首次实现跨数据集、跨关键点的全身体态估计,性能显著优于现有方法。
  • 支持深度信息输入,适合需要高精度姿态建模的研究与应用。

在计算机视觉快速发展的背景下,从多视角准确估计多人姿态是一项严峻挑战,尤其要求结果具备可靠性。本文系统评估了多视角多人姿态估计算法在未见数据集上的泛化能力,并提出一种新算法,在该任务中表现优异。研究还探讨了引入深度信息带来的性能提升。所提方法不仅能有效泛化至未见数据集,还能适应不同关键点,首次实现了多视角多人全身体态估计。为促进后续研究,所有成果均公开可用。

原文摘要 · Abstract (English)

In the rapidly evolving field of computer vision, the task of accurately estimating the poses of multiple individuals from various viewpoints presents a formidable challenge, especially if the estimations should be reliable as well. This work presents an extensive evaluation of the generalization capabilities of multi-view multi-person pose estimators to unseen datasets and presents a new algorithm with strong performance in this task. It also studies the improvements by additionally using depth information. Since the new approach can not only generalize well to unseen datasets, but also to different keypoints, the first multi-view multi-person whole-body estimator is presented. To support further research on those topics, all of the work is publicly accessible.

姿态估计多视角泛化能力体素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。