对比人脸与人体追踪,提升人机交互中的身份连续性。
Face versus Body Tracking for Human-Robot Interaction: An Egocentric Dataset

- 区分检测误差与追踪逻辑,系统评估人脸与人体追踪效果。
- 引入重识别机制使身体追踪更稳定,但人脸追踪误识别率上升49%。
- 适用于社交机器人场景,尤其关注近距离遮挡与动态交互。
有意义的人机交互(HRI)要求机器人通过持续追踪用户来评估其参与度。然而,当前主流多目标追踪模型主要针对监控或自动驾驶优化,难以应对社交机器人面临的独特视角挑战:人类移动路径非线性、相互遮挡或进出画面,易引发频繁的身份切换(IDSW),导致对话中断。为此,我们基于Furhat机器人构建了一个专注的自视点数据集,并系统评估了检测误差与追踪逻辑的影响,比较了人脸与人体追踪策略,分析了延长记忆和外观重识别(ReID)的作用。结果表明,增加时间记忆可缓解长时间遮挡,但在复杂动态事件中仍无效;整合ReID能显著提升身体追踪稳定性,却因侧脸角度敏感导致人脸身份切换激增。最终,优化后的流程相比标准追踪基线将身份切换减少49%,有效缓解交互中断。由于现有基准缺乏密集近距离遮挡,本工作强调必须采用原生采集的社会动态数据才能真正验证HRI感知模型。
原文摘要 · Abstract (English)
Meaningful human-robot interaction (HRI) requires a robot to continuously assess user engagement through persistent user tracking. However, state-of-the-art Multi-Object Tracking models are heavily optimized for surveillance or autonomous driving. A social robot faces distinct egocentric challenges, such as humans moving in unpredictable nonlinear patterns, obstructing each other, or leaving and reentering the scene. These dynamics trigger frequent identity switches (IDSW), causing the robot to lose its footing mid-conversation. To address this, we introduce a focused, custom-annotated egocentric dataset collected via the Furhat robot. We present a systematic evaluation isolating detection errors from tracking logic, comparing face versus body tracking, and assessing the impact of extended memory and appearance re-identification (ReID). Results indicate that increasing temporal memory mitigates prolonged occlusions but fails on complex dynamic events. Integrating ReID resolves complex switches but exhibits opposing effects: it substantially improves body tracking stability, yet causes facial IDSW to spike due to profile angle sensitivity. Ultimately, our optimized pipeline reduces IDSW by 49% compared to a standard tracking-by-detection baseline, effectively mitigating interaction breakdowns. As standard benchmarks lack dense, close-quarter occlusions, this work highlights the critical need for natively captured social dynamics to truly validate HRI perception models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。