arXiv:2609.07091cs.RO2026-09

机器人遇人被挡时,靠地图推理下一步去哪,找回目标成功率超71%。

Human-Aware Target Tracking and Navigation: Fusing Kinematic State Estimation with Structural Map Constraints

论文配图:Human-Aware Target Tracking and Navigation: Fusing Kinematic State Estimation with Structural Map Constraints
图 1 · 摘自论文原文
  • 融合视觉、激光与地图,实时推断人的可能行进路线。
  • 7秒内遮挡后仍能成功找回目标,成功率71.4%。
  • 适合复杂人流场景下的智能跟随,如医院或商场。

在动态环境中,自主移动机器人在追踪人类目标时常因遮挡或传感器丢失导致跟踪失败。本文提出一个端到端的自主导航系统,通过地图引导的空间推理解决目标遮挡问题。系统采用多模态感知管道,融合基于深度学习的视觉跟踪与二维激光雷达点云聚类,实现对目标人物的高保真跟踪。连续状态估计器整合感知数据、轮式里程计和惯性测量单元(IMU)以实现稳定定位。当主动跟踪因遮挡丢失时,系统启动基于地图的恢复框架:利用预设拓扑地图,执行图搜索,沿结构化步行路径传播目标最后已知轨迹,并遵循左侧行走区域规范。系统生成一系列可行的未来轨迹,推理可能的结构变化(如直行或路口转弯),并将这些预测直接输入局部避障规划器,使机器人能在重新视觉捕获前安全、可预测地持续跟随。真实世界测试在密集多人环境下验证了系统的鲁棒性,面对长达7秒的主要遮挡事件,目标重获成功率达到71.4%。

原文摘要 · Abstract (English)

Autonomous mobile robots performing person-following tasks often suffer from temporary occlusions and sensor track loss in dynamic environments. This research presents an end-to-end autonomous navigation stack that addresses target occlusion through map-informed spatial reasoning. The proposed system features a multi-modal perception pipeline, fusing deep learning-based visual tracking with 2-dimensional LiDAR point clustering to maintain high-fidelity tracking of a tagged person. A continuous state estimator integrates this perception data with wheel odometry and IMU sensors for stable localization. When the active track is lost due to occlusion, the system activates a map-based recovery framework. Leveraging a predefined topological map, the system executes a graph-based search to propagate the target's last known trajectory along structurally defined walking lanes, adhering to left-hand regional conventions. By generating a discrete set of feasible future trajectories, the robot reasons about potential structural trajectory changes, such as continuing a heading or turning at an intersection. This map-informed prediction is fed directly to the local obstacle avoidance planner, enabling the robot to continue following its target safely and predictably until the person is visually reacquired. Real-world evaluations in dense multi-person environments demonstrate the system's robustness, achieving a 71.4\% target reacquisition success rate during major occlusion events lasting up to 7 seconds.

目标追踪地图推理机器人导航多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。