让视觉语言模型学会空间推理和长期记忆,提升机器人追踪动态目标能力
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- 引入极坐标思维链,实时推断目标相对位置
- 在严重遮挡下仍保持追踪,比之前方法提升12分
- 适合需要长时间稳定追踪的机器人应用
具身视觉追踪(EVT)是陪伴机器人、导览机器人和服务助手等实际应用的核心能力,要求持续跟踪移动目标。现有方法缺乏显式空间推理与有效时序记忆,在严重遮挡或存在相似干扰物时易失败。为此,我们提出TrackVLA++,一种新型视觉-语言-动作(VLA)模型,包含两个关键模块:空间推理机制与目标识别记忆(TIM)。推理模块采用名为Polar-CoT的思维链范式,推断目标相对位置并编码为紧凑极坐标标记用于动作预测。基于这些空间先验,TIM采用门控更新策略,实现长时程目标记忆,保障时空一致性,缓解长时间遮挡导致的目标丢失。大量实验表明,TrackVLA++在公开基准上取得当前最优性能,尤其在EVT-Bench DT数据集上分别超越先前领先方法5.1和12分。此外,该模型展现出强零样本泛化能力,可在动态且遮挡复杂的现实场景中实现鲁棒追踪。
原文摘要 · Abstract (English)
Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have enabled language-guided tracking in complex and unstructured scenes. However, existing approaches lack explicit spatial reasoning and effective temporal memory, causing failures under severe occlusions or in the presence of similar-looking distractors. To address these challenges, we present TrackVLA++, a novel Vision-Language-Action (VLA) model that enhances embodied visual tracking with two key modules, a spatial reasoning mechanism and a Target Identification Memory (TIM). The reasoning module introduces a Chain-of-Thought paradigm, termed Polar-CoT, which infers the target's relative position and encodes it as a compact polar-coordinate token for action prediction. Guided by these spatial priors, the TIM employs a gated update strategy to preserve long-horizon target memory, ensuring spatiotemporal consistency and mitigating target loss during extended occlusions. Extensive experiments show that TrackVLA++ achieves state-of-the-art performance on public benchmarks across both egocentric and multi-camera settings. On the challenging EVT-Bench DT split, TrackVLA++ surpasses the previous leading approach by 5.1 and 12, respectively. Furthermore, TrackVLA++ exhibits strong zero-shot generalization, enabling robust real-world tracking in dynamic and occluded scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。