arXiv:2503.02916cs.CVcs.RO2025-03中稿 · IROS2025被引 2

基于四点人体模型,联合估计相机姿态与3D位置,提升运动相机下人物定位精度。

Monocular Person Localization under Camera Ego-motion

  • 将人体建模为四点,通过优化联合求解2D相机姿态与3D人体位置。
  • 在公开数据集与真实机器人实验中,定位误差显著低于基线方法。
  • 适用于移动相机下的机器人跟随系统,已部署于敏捷四足机器人。

从运动单目相机中定位人物对人机交互至关重要。现有方法或依赖固定相机几何假设,或在缺乏相机自运动的数据集上训练位置回归模型,难以应对严重相机自运动,导致定位不准。本文将人物定位视为姿态估计问题,采用四点人体模型,通过优化联合估计2D相机姿态与人物3D位置。在多个公开数据集及真实机器人实验中,本方法均优于基线,定位精度显著提升。该方法已集成至人物跟随系统,并成功部署于敏捷四足机器人平台。

原文摘要 · Abstract (English)

Localizing a person from a moving monocular camera is critical for Human-Robot Interaction (HRI). To estimate the 3D human position from a 2D image, existing methods either depend on the geometric assumption of a fixed camera or use a position regression model trained on datasets containing little camera ego-motion. These methods are vulnerable to severe camera ego-motion, resulting in inaccurate person localization. We consider person localization as a part of a pose estimation problem. By representing a human with a four-point model, our method jointly estimates the 2D camera attitude and the person's 3D location through optimization. Evaluations on both public datasets and real robot experiments demonstrate our method outperforms baselines in person localization accuracy. Our method is further implemented into a person-following system and deployed on an agile quadruped robot.

单目定位相机自运动人机交互四足机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。