arXiv:2508.03541cs.ROcs.LG2025-08

用单目摄像头实现多行人追踪与姿态估计,提升机器人避障能力。

Vision-based Perception System for Automated Delivery Robot-Pedestrians Interactions

  • 基于单个摄像头完成多人检测、追踪与姿态估计
  • 在密集人群下身份保持率提升10%,追踪准确率达85%以上
  • 可识别弱势群体,让机器人行为更符合社会规范

自动化配送机器人(ADRs)进入人流密集的城市空间,带来安全、高效且符合社交规范导航的挑战。本文构建了一套基于单个视觉传感器的完整系统,实现多行人检测与追踪、姿态估计及单目深度感知。利用真实世界MOT17数据集序列,研究证明融合人体姿态估计与深度信息能显著提升行人轨迹预测与身份维持能力,即使在遮挡和高密度人群场景下表现优异。结果表明,系统在身份保持率(IDF1)上最高提升10%,多目标追踪准确率(MOTA)提升7%,检测精度始终超过85%。此外,系统能识别弱势行人群体,支持更具社会意识与包容性的机器人行为。

原文摘要 · Abstract (English)

The integration of Automated Delivery Robots (ADRs) into pedestrian-heavy urban spaces introduces unique challenges in terms of safe, efficient, and socially acceptable navigation. We develop the complete pipeline for a single vision sensor based multi-pedestrian detection and tracking, pose estimation, and monocular depth perception. Leveraging the real-world MOT17 dataset sequences, this study demonstrates how integrating human-pose estimation and depth cues enhances pedestrian trajectory prediction and identity maintenance, even under occlusions and dense crowds. Results show measurable improvements, including up to a 10% increase in identity preservation (IDF1), a 7% improvement in multiobject tracking accuracy (MOTA), and consistently high detection precision exceeding 85%, even in challenging scenarios. Notably, the system identifies vulnerable pedestrian groups supporting more socially aware and inclusive robot behaviour.

视觉感知机器人导航行人追踪单目深度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。