arXiv:2506.19747cs.CVcs.RO2025-06ICRA被引 2

对比多种投影方法,提升鱼眼图像中3D人体姿态估计精度。

Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images

  • 用针孔、等距和双球模型等投影法处理鱼眼畸变。
  • 双球模型显著提升广角下人体姿态估计准确率。
  • 提出基于检测框的投影模型选择策略,适合机器人交互场景。

鱼眼相机相比标准针孔相机具有更宽的视场(FOV),在人机交互和汽车应用中尤为有用。然而,由于鱼眼光学固有的弯曲畸变,从单目鱼眼图像中准确检测人体姿态仍具挑战性。尽管已有多种鱼眼图像去畸变方法,但其在覆盖广角的人体姿态估计中的有效性与局限性尚未系统评估。为此,我们评估了针孔、等距和双球相机模型以及圆柱投影方法对3D人体姿态估计精度的影响。结果表明,在近距离场景中,针孔投影不适用,最优投影方法随姿态覆盖的视场变化而变化。使用双球等先进鱼眼模型可显著提升3D人体姿态估计精度。我们提出一种基于检测边界框的投影模型选择启发式方法以提升预测质量。此外,我们构建并评估了新数据集FISHnCHIPS,包含鱼眼图像中3D人体骨骼标注,涵盖极端近距、地面安装相机及大视场姿态等非常规视角,数据集地址:https://www.vision.rwth-aachen.de/fishnchips

原文摘要 · Abstract (English)

Fisheye cameras offer robots the ability to capture human movements across a wider field of view (FOV) than standard pinhole cameras, making them particularly useful for applications in human-robot interaction and automotive contexts. However, accurately detecting human poses in fisheye images is challenging due to the curved distortions inherent to fisheye optics. While various methods for undistorting fisheye images have been proposed, their effectiveness and limitations for poses that cover a wide FOV has not been systematically evaluated in the context of absolute human pose estimation from monocular fisheye images. To address this gap, we evaluate the impact of pinhole, equidistant and double sphere camera models, as well as cylindrical projection methods, on 3D human pose estimation accuracy. We find that in close-up scenarios, pinhole projection is inadequate, and the optimal projection method varies with the FOV covered by the human pose. The usage of advanced fisheye models like the double sphere model significantly enhances 3D human pose estimation accuracy. We propose a heuristic for selecting the appropriate projection model based on the detection bounding box to enhance prediction quality. Additionally, we introduce and evaluate on our novel dataset FISHnCHIPS, which features 3D human skeleton annotations in fisheye images, including images from unconventional angles, such as extreme close-ups, ground-mounted cameras, and wide-FOV poses, available at: https://www.vision.rwth-aachen.de/fishnchips

3D姿态估计鱼眼图像投影模型机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。