arXiv:2509.17141cs.RO2025-09中稿 · ICRA被引 14

通过点追踪构建物体中心的历史表征,提升机器人操作的长期记忆能力。

History-Aware Visuomotor Policy Learning via Point Tracking

  • 用点追踪抽象历史观察,生成紧凑且结构化的物体级记忆
  • 在多种操作任务中显著提升任务成功率与决策准确率
  • 适合需要长时序记忆的复杂机器人控制场景

许多操作任务需要超越当前观测的记忆能力,但现有视觉-运动策略多依赖马尔可夫假设,在重复状态或长时程依赖下表现不佳。现有方法虽尝试扩展观测时长,仍难以满足多样化的记忆需求。为此,我们提出基于点追踪的物体中心历史表征,将过往观测抽象为仅保留关键任务信息的紧凑结构化形式。追踪点在物体层级编码并聚合,生成可无缝集成至各类视觉-运动策略的紧凑历史表示。该设计实现完整的历史感知与高计算效率,显著提升整体任务表现与决策准确性。在多种操作任务上的广泛评估表明,该方法有效应对任务阶段识别、空间记忆、动作计数等多重记忆需求,以及连续与预加载记忆等长期需求,并持续优于马尔可夫基线及已有历史增强方法。

原文摘要 · Abstract (English)

Many manipulation tasks require memory beyond the current observation, yet most visuomotor policies rely on the Markov assumption and thus struggle with repeated states or long-horizon dependencies. Existing methods attempt to extend observation horizons but remain insufficient for diverse memory requirements. To this end, we propose an object-centric history representation based on point tracking, which abstracts past observations into a compact and structured form that retains only essential task-relevant information. Tracked points are encoded and aggregated at the object level, yielding a compact history representation that can be seamlessly integrated into various visuomotor policies. Our design provides full history-awareness with high computational efficiency, leading to improved overall task performance and decision accuracy. Through extensive evaluations on diverse manipulation tasks, we show that our method addresses multiple facets of memory requirements - such as task stage identification, spatial memorization, and action counting, as well as longer-term demands like continuous and pre-loaded memory - and consistently outperforms both Markovian baselines and prior history-based approaches. Project website: http://tonyfang.net/history

机器人控制历史记忆点追踪视觉运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。