arXiv:2505.07446cs.RO2025-05被引 6

构建大规模机器人视角目标人物追踪数据集,支持复杂环境长期跟踪研究。

TPT-Bench: A Large-Scale, Long-Term and Robot-Egocentric Dataset for Benchmarking Target Person Tracking

  • 通过机器人跟随真人采集多模态数据,模拟真实场景下的人机交互。
  • 包含48段序列、长时遮挡与频繁重识别挑战,覆盖室内外复杂环境。
  • 适合研究长时目标追踪、人机协作及具身智能的学者使用。

从机器人第一视角追踪目标人物对于实现自主机器人在人机交互与具身智能中提供持续个性化协助至关重要。然而,现有目标人物追踪(TPT)基准大多局限于受控实验室环境,存在干扰少、背景干净、遮挡时间短等问题。本文提出一个大规模数据集,用于在拥挤且非结构化环境中评估目标人物追踪。数据由真人推动搭载传感器的小车跟随目标人物采集,体现类人跟随行为,突出长期追踪挑战,包括频繁遮挡和从众多行人中重新识别的需求。数据集包含里程计、3D LiDAR、IMU、全景图像和RGB-D图像等多模态数据流,并对48个序列中的目标人物进行了详尽的2D边界框标注,涵盖室内外场景。基于该数据集与视觉标注,我们对现有SOTA TPT方法进行了广泛实验,深入分析其局限性,并为未来研究指明方向。

原文摘要 · Abstract (English)

Tracking a target person from robot-egocentric views is crucial for developing autonomous robots that provide continuous personalized assistance or collaboration in Human-Robot Interaction (HRI) and Embodied AI. However, most existing target person tracking (TPT) benchmarks are limited to controlled laboratory environments with few distractions, clean backgrounds, and short-term occlusions. In this paper, we introduce a large-scale dataset designed for TPT in crowded and unstructured environments, demonstrated through a robot-person following task. The dataset is collected by a human pushing a sensor-equipped cart while following a target person, capturing human-like following behavior and emphasizing long-term tracking challenges, including frequent occlusions and the need for re-identification from numerous pedestrians. It includes multi-modal data streams, including odometry, 3D LiDAR, IMU, panoramic images, and RGB-D images, along with exhaustively annotated 2D bounding boxes of the target person across 48 sequences, both indoors and outdoors. Using this dataset and visual annotations, we perform extensive experiments with existing SOTA TPT methods, offering a thorough analysis of their limitations and suggesting future research directions.

目标追踪机器人视觉多模态数据具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。