arXiv:2606.25880cs.CV2026-06

用空间+语义提示提升机器人视觉追踪精度

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning

论文配图:USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning
图 1 · 摘自论文原文
  • 融合文本、点、框、掩码的统一提示框架
  • 空间提示使追踪成功率显著提升,尤其在长时追踪中
  • 支持实时推理,适合真实机器人部署

具身视觉追踪(EVT)要求智能体在动态环境中主动移动并持续跟踪目标。现有方法多依赖语言描述,但在复杂场景中常因多个物体满足相同语义而造成目标混淆。为此,我们提出从纯文本提示转向统一的空间-语义提示的新范式。基于此,我们构建了USS框架,可在统一架构中支持文本、点、边界框和掩码四种提示形式。USS使用模态专用编码器处理异构提示,通过混合注意力将提示与视觉特征融合,并解码出以提示为条件的紧凑表征,生成以我为中心的导航点。为进一步增强时间鲁棒性,引入潜在世界模型,通过自监督对齐预测未来表示。真实机器人实验表明,显式空间提示在含相似干扰物和长时追踪场景中显著提升成功率。仿真基准测试中,USS在非多模态大模型方法中达到领先性能,且优于部分多模态大模型方法,推理速度更快。结果表明,空间-语义提示为具身视觉追踪提供了更精确灵活的目标指示方式。

原文摘要 · Abstract (English)

Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However, prevailing EVT paradigms predominantly rely on language-based target indication. While language is expressive and convenient, cluttered scenes often contain multiple objects that satisfy the same semantic description, leading to ambiguous target grounding. We therefore propose a paradigm shift, reframing target indication in EVT from text-only specification to unified spatial-semantic prompting. Based on this paradigm, we introduce Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning, USS, an end-to-end embodied tracking framework that supports text, point, bounding box, and mask prompts within a unified architecture. USS encodes heterogeneous prompts with modality-specific encoders, fuses prompt tokens with visual features through hybrid attention, and decodes compact prompt-conditioned representations into egocentric waypoints. To further improve temporal robustness, USS incorporates a latent world model that predicts future representations through self-supervised alignment. Real-robot experiments demonstrate that explicit spatial target cues yield higher success rates than text-only prompts, particularly in scenarios involving similar distractors and longer-horizon tracking where maintaining instance-level target identity is critical. In the simulation benchmark, USS also achieves state-of-the-art performance among non-MLLM-based methods and competitive results against recent MLLM-based approaches with faster inference speed. Our findings reveal that spatial-semantic prompting provides a more precise and flexible target indication interface for embodied visual tracking. Project site: https://arescheah.github.io/uss-project-page/.

视觉追踪具身智能空间提示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。