arXiv:2505.20718cs.CVcs.AI2025-05被引 7

用视觉语言模型让机器人跟踪失败后自动恢复,提升追踪成功率。

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

  • 跟踪正常时用快速策略,失败时调用语言模型推理修复
  • 通过记忆反思机制,模型能从过往经验中持续改进
  • 在复杂环境里成功率提升220%,适合需要持续监控的机器人

我们提出一种新型自提升框架,利用视觉语言模型(VLM)增强具身视觉追踪(EVT),解决现有主动追踪系统在跟踪失败后难以恢复的问题。该方法将现成的主动追踪技术与VLM的推理能力结合,日常追踪采用快速视觉策略,仅在检测到失败时激活VLM进行推理。框架引入基于记忆的自我反思机制,使VLM能从过往经验中逐步提升,有效克服其在三维空间推理上的不足。实验表明,该框架在挑战性环境中,相较最先进的强化学习方法成功率提升72%,相较PID方法提升220%。这是首个将基于VLM的推理用于主动协助EVT代理实现前瞻性失败恢复的工作,为需要在动态非结构化环境中持续监控目标的现实机器人应用带来显著进展。项目网站:https://sites.google.com/view/evt-recovery-assistant。

原文摘要 · Abstract (English)

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with VLMs' reasoning capabilities, deploying a fast visual policy for normal tracking and activating VLM reasoning only upon failure detection. The framework features a memory-augmented self-reflection mechanism that enables the VLM to progressively improve by learning from past experiences, effectively addressing VLMs' limitations in 3D spatial reasoning. Experimental results demonstrate significant performance improvements, with our framework boosting success rates by $72\%$ with state-of-the-art RL-based approaches and $220\%$ with PID-based methods in challenging environments. This work represents the first integration of VLM-based reasoning to assist EVT agents in proactive failure recovery, offering substantial advances for real-world robotic applications that require continuous target monitoring in dynamic, unstructured environments. Project website: https://sites.google.com/view/evt-recovery-assistant.

视觉追踪语言模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。