通过功能物体归一化,让机器人学会可复用的通用操作动作。
FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
- 将长时序操作拆解为动作块,聚焦动作本身而非具体任务。
- 利用视觉语言模型提取物体功能线索,实现跨物体轨迹迁移。
- 支持类别泛化、跨任务复用,适合复杂操作场景的仿生学习。
端到端演示生成的通用机器人技能往往导致仅在训练分布内有效的任务特定策略。为此,我们提出 FunCanon 框架,将长时序操作任务转化为由执行者、动词和物体定义的动作块序列。该方法聚焦于动作本身,而非孤立任务,从而实现组合性与复用性。为使策略具备位姿感知与类别泛化能力,我们基于大视觉语言模型提供的功能线索,对物体进行功能归一化,实现功能对齐与自动轨迹迁移,将物体映射至共享功能坐标系。在此对齐数据上训练的以物体为中心和动作为中心的扩散策略 FuncDiffuser,自然尊重物体功能与位姿,简化学习过程并提升泛化能力。在模拟与真实世界基准测试中,实验表明该方法具备类别级泛化、跨任务行为复用及鲁棒的 sim2real 部署能力,证明功能归一化为复杂操作领域中的可扩展模仿学习提供了强大归纳偏置。
原文摘要 · Abstract (English)
General-purpose robotic skills from end-to-end demonstrations often leads to task-specific policies that fail to generalize beyond the training distribution. Therefore, we introduce FunCanon, a framework that converts long-horizon manipulation tasks into sequences of action chunks, each defined by an actor, verb, and object. These chunks focus policy learning on the actions themselves, rather than isolated tasks, enabling compositionality and reuse. To make policies pose-aware and category-general, we perform functional object canonicalization for functional alignment and automatic manipulation trajectory transfer, mapping objects into shared functional frames using affordance cues from large vision language models. An object centric and action centric diffusion policy FuncDiffuser trained on this aligned data naturally respects object affordances and poses, simplifying learning and improving generalization ability. Experiments on simulated and real-world benchmarks demonstrate category-level generalization, cross-task behavior reuse, and robust sim2real deployment, showing that functional canonicalization provides a strong inductive bias for scalable imitation learning in complex manipulation domains. Details of the demo and supplemental material are available on our project website https://sites.google.com/view/funcanon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。