arXiv:2606.07233cs.CVcs.LG2026-06中稿 · publication at the…被引 1

轻量级视觉追踪提升机器人3D行人跟踪连续性

Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking

论文配图:Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking
图 1 · 摘自论文原文
  • 用投影框架分离几何与外观建模,降低计算开销
  • 级联匹配策略有效恢复遮挡轨迹,避免身份错乱
  • 轻量CNN与ViT适配移动机器人,兼顾速度与精度

基于激光雷达的3D多目标跟踪通常仅依赖几何信息,在长时间遮挡或人群密集场景中难以区分目标。虽引入基于RGB的重识别(ReID)可维持身份上下文,但现有方法常需计算量大的并行检测器,影响机器人实时响应。本文系统研究图像驱动的ReID在在线3D多行人跟踪中的应用,提出轻量级投影框架,解耦几何与外观建模,适用于移动机器人。通过对比轻量级CNN与视觉变换器(Vision Transformers)的特征提取架构,并评估多种多模态数据关联策略,平衡延迟与跟踪鲁棒性。在KITTI数据集行人类别上的实验表明,简单的线性融合会因视觉噪声导致性能下降;而级联匹配策略能有效恢复遮挡轨迹,不牺牲整体精度,防止身份切换,保障人机交互连贯性。结果证明,轻量级模型可在低延迟导航需求与社会感知所需的判别能力间实现最优权衡。

原文摘要 · Abstract (English)

LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments. While integrating RGB-based Re-Identification (ReID) offers a theoretical solution for preserving identity context, existing approaches often rely on computationally expensive parallel detectors that hinder real-time robot responsiveness. This work presents a systematic study of image-based ReID in online 3D MOT, utilizing a lightweight projection-based framework to decouple geometric and appearance modeling for mobile robots. A comprehensive analysis of feature extraction architectures is conducted, employing lightweight CNNs and Vision Transformers, and evaluating various multi-modal data association strategies to balance computational latency with robust tracking. Experiments on the Pedestrian class of the KITTI dataset reveal that naive linear fusion, of appearance and motion costs, degrades performance due to visual noise. Conversely, a cascaded matching strategy successfully recovers occluded tracks without compromising overall precision, effectively preventing identity switches to maintain human-robot interaction continuity. We show that lightweight architectures can offer an optimal trade-off between the low latency required for safe navigation and the discriminative power needed for social awareness.

3D跟踪视觉重识别轻量模型机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。