从人类示范中学习语义-几何联合表示,提升双臂操作的机器人规划能力。
Semantic-Geometric Task Representations for Bimanual Manipulation from Human Demonstrations to Robot Action Planning
- 用图神经网络与Transformer构建语义-几何联合表征,解耦动作标签
- 在11个任务上提升规划成功率,尤其在动作顺序多变时优势明显
- 适合需要灵活适应复杂双臂操作的机器人研究者使用
从人类示范中学习结构化任务表征对双臂操作至关重要,因动作顺序、物体参与及交互几何关系在不同执行中差异显著。核心挑战在于同时捕捉离散语义任务结构与以物体为中心的时空几何关系演化,以支持任务进展推理。本文提出一种基于语义-几何图的任务表征方法,通过消息传递神经网络(MPNN)编码器与Transformer解码器联合编码物体身份、物体间语义关系及每物体运动历史。编码器仅作用于时序场景图,生成与动作标签解耦的结构化表示;解码器则结合动作上下文预测未来动作、关联物体与物体运动。该解耦机制实现任务无关表示,可仅通过微调解码器在少量机器人数据上复用编码器。在来自两个数据集的11个双臂任务中,结构化语义-几何表征相比简单序列模型的优势随动作顺序和物体参与度变化而增大。部署时,规划器将动作与运动预测结合学习到的概率运动基元,在两个真实机器人双臂任务中实现全任务成功,并优于图结构消融实验、Transformer、仅解码器模型及微调视觉语言模型基线。
原文摘要 · Abstract (English)
Learning structured task representations from human demonstrations is essential for bimanual manipulation, where action ordering, object involvement, and interaction geometry vary significantly across executions. A key challenge lies in jointly capturing the discrete semantic task structure and the temporal evolution of object-centric geometric relations in a form that supports reasoning over task progression. We introduce a semantic--geometric graph-based task representation that jointly encodes object identities, inter-object semantic relations, and per-object motion histories, via a Message Passing Neural Network (MPNN) encoder and a Transformer-based decoder. The encoder operates solely on the temporal scene graph, producing structured representations decoupled from action labels. The decoder then conditions on action-context to forecast future actions, associated objects, and object motions. This decoupling learns task-agnostic representations, enabling encoder reuse across embodiments through decoder-only finetuning on a small robot dataset. Across eleven bimanual tasks from two datasets, we find that the benefit of structured semantic--geometric representations over simpler sequence-based models grows with task variability in action ordering and object involvement. At deployment, a planner couples the action and motion predictions with learned Probabilistic Movement Primitives, achieving full task success on two real-robot bimanual tasks and outperforming graph ablations, Transformer, decoder-only, and finetuned vision-language model baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。