聚焦物体间交互,让机器人更高效学习复杂操作。
FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation Policies

- 从示范中自动提取短时交互片段,实现时空双重抽象
- 仅用少量数据即达领先性能,显著降低训练成本
- 适合需要泛化能力的物理操作任务,如抓取、装配
以物体为中心的操作策略通过建模物体运动而非直接预测机器人动作,提升了泛化能力。然而,现有方法常受限于表示形式:过于简单难以捕捉交互动态,或过于密集导致学习效率低。我们观察到,许多刚性关系操作任务由短暂的交互阶段主导,此时关键物体间的相对运动高度受约束。基于此,我们提出Foci Policy,一种以交互为中心的框架,实现双重抽象:(1) 时间上,自动从示范中提取紧凑的交互片段;(2) 空间上,将技能表示为任务相关物体间的相对SE(3)运动,实现对场景配置和机器人本体的不变性。在RLBench、COLOSSEUM及真实世界任务上的实验表明,Foci Policy以远少于以往物体中心与动作中心策略的训练数据,取得了优异表现。结果表明,建模物体间交互为刚性关系操作提供了简单而高效的归纳偏置。
原文摘要 · Abstract (English)
Object-centric manipulation policies improve generalization by modeling object motion instead of directly predicting robot actions. However, existing methods are often limited by representations which are either too simplistic to capture interaction dynamics or too dense to learn efficiently. We observe that many rigid relational manipulation tasks are governed by short interaction phases where the relative motion between task-relevant objects is tightly constrained. Based on this observation, we propose \textsc{Foci Policy}, an interaction-centric framework that achieves a two-fold abstraction: (1) temporally, by automatically extracting compact interaction segments from demonstrations;(2) spatially, by representing skills as relative $SE(3)$ motion between task-relevant objects, yielding invariance to scene configurations and robot embodiment. Experiments on RLBench, COLOSSEUM, and real-world tasks show that \textsc{Foci Policy} achieves strong performance with substantially less training data than prior object-centric and action-centric policies. These results suggest that modeling object-object interactions provides a simple and efficient inductive bias for rigid relational manipulation. Project page: \href{https://fitz0401.github.io/foci-page/}{fitz0401.github.io/foci-page/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。