arXiv:2609.06302cs.CVcs.RO2026-09

提出因果结构世界模型,解决机器人追踪目标时的错误预测问题。

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

论文配图:CST-WM: A Causally Structured World Model for Embodied Visual Tracking
图 1 · 摘自论文原文
  • 将潜空间分解为目标、机器人和观测三部分,阻断动作对目标证据的直接干扰
  • 在真实场景中实现稳定跟踪与丢失后重捕,性能超越传统基线模型
  • 适合需要长期规划的机器人视觉追踪任务,尤其关注鲁棒性与可解释性

具身视觉追踪要求机器人不仅响应当前视图,还需选择动作以维持或恢复因自身运动、遮挡和干扰导致的目标可见性。这本质上是一个关于未来目标可观测性和表观尺度的预测决策问题。核心难点在于一种任务特异的因果幻觉:在动作条件预测中,模型可能利用机器人控制与目标观测之间的强相关性,错误地幻想动作对目标证据有直接因果影响,而非通过机器人运动带来的观测变化间接影响。这种捷径虽生成看似合理的未来,但语义错误,不利于追踪规划与重捕。我们提出CST-WM,一种因果结构化世界模型,将潜状态分解为目标证据、机器人和观测分支,并因子化转移机制,阻止动作直接注入目标证据分支,同时保留动作对机器人运动和观测更新的影响。结合基于回溯的模型预测控制,CST-WM可在同一框架下支持稳定跟随与临时目标重捕。在EVT-Bench和Habitat 3.0上,覆盖标准追踪、目标丢失恢复与跨数据集迁移,其在跟踪质量、距离范围控制、安全性及重捕能力方面均优于反应式与世界模型基线;离线诊断显示其多步回溯保真度更高,规划价值一致性更强,且显著减少动作直接泄漏。

原文摘要 · Abstract (English)

Embodied visual tracking requires a robot not only to react to the current view, but to choose actions that preserve or recover future evidence of a moving target under ego-motion, occlusion, and distractors. It is therefore a predictive decision problem over future target observability and apparent scale. A central difficulty is a task-specific form of causal hallucination: in action-conditioned prediction, a model can exploit the strong correlation between robot control and target-related observations by hallucinating a direct causal effect from the current action to target evidence, rather than letting action influence that evidence only through robot motion and the resulting observation change. The shortcut yields plausible futures with the wrong semantics for tracking-oriented planning and re-acquisition. We propose CST-WM, a causally structured world model that decomposes the latent state into target-evidence, robot, and observation branches and factorizes the transition so that direct action injection into the target-evidence branch is blocked, while action remains available to robot motion and observation updates. Combined with rollout-based model-predictive control, CST-WM supports both stable following and temporary target re-acquisition in one planning framework. On EVT-Bench and Habitat 3.0, covering standard tracking, target-loss recovery, and cross-dataset transfer, it improves following quality, distance-range control, safety, and re-acquisition over reactive and world-model baselines; offline diagnostics show better multi-step rollout fidelity, stronger planning-value consistency, and substantially reduced direct action leakage. For embodied visual tracking, future prediction alone is not enough: the predictive structure itself must align with how target evidence enters planning.

视觉追踪因果建模具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。