arXiv:2607.00031cs.RO2026-07

通过预测交互效果,让机器人自动发现物体和动作的抽象符号。

Joint Discovery of Object and Action Symbols through Effect Prediction for Robotic Manipulation Planning

论文配图:Joint Discovery of Object and Action Symbols through Effect Prediction for Robotic Manipulation Planning
图 1 · 摘自论文原文
  • 用二元瓶颈层从随机交互中联合学习物体与动作符号。
  • 在桌面上重排和堆叠任务中,对新物体实现少样本泛化,成功率超基准方法。
  • 适合需要理解行为差异而非外观相似性的机器人操作场景。

为实现复杂操作规划,自主机器人需将连续、高维的感知运动交互抽象为离散的物体与动作表示。以往方法或基于视觉外观分类物体,无法区分外观相似但行为不同的物体;或基于交互结果分类,但仅限预定义动作。为此,我们提出一种模型,通过二元瓶颈层联合发现高层操作原语与物体类别,训练目标是预测多模态结果(包括物体运动、接触状态和力反馈),输入为随机交互数据。基于这些发现的二元表示,采用离散规划方法,利用预测效果轨迹中的中间步骤,实现部分动作执行以达成精确底层控制。此外,我们通过少量交互效果与学习到的物体符号进行对比,对新物体进行类别分配,实现基于行为而非外观的少样本泛化。我们在桌面重排与堆叠任务上进行实验,结果表明,该效果驱动的规划方法在已见与新物体上的规划精度均优于当前最优方法与基于视觉的替代方案。

原文摘要 · Abstract (English)

To perform complex manipulation planning, autonomous robots are required to abstract continuous, high-dimensional sensorimotor interactions into discrete object and action representations. Earlier work either categorized objects based on visual appearances, which fails to distinguish objects that appear similar but behave differently, or based on effects under interaction, but was limited to predefined actions. To address these limitations, we propose a model that jointly discovers high-level manipulation primitives and object categories through a binary bottleneck layer, trained to predict multi-modal outcomes, including object motion, contact, and force feedback, from random interaction data. Building on these discovered binary representations, we leverage a discrete planning method that uses intermediate steps in the predicted effect trajectory to enable partial action executions for precise low-level control. Additionally, we evaluate our framework's generalization capabilities on novel objects by assigning object categories through comparing a small number of interaction effects with the predicted effects of learned object symbols, enabling few-shot generalization based on behavior rather than visual similarity. We conduct experiments on tabletop repositioning and stacking tasks, and confirm that our effect-driven planning approach outperforms both a state-of-the-art method and a visual-based alternative in planning precision across seen and novel objects.

机器人操作符号学习少样本泛化效果预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。