arXiv:2410.14974cs.RO2024-10ICRA被引 25

用因果注意力提升机器人在少样本下的泛化能力

CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation

  • 引入因果注意力机制,结合DINOv2与LoRA增强环境理解
  • 仅需50次演示即在多变场景中实现43%任务完成率
  • 适合少样本、跨环境部署的机器人操作研究者

机器人操作的泛化仍是关键挑战,尤其在新环境中示范数据有限时。本文提出CAGE,一种新型机器人操作策略,通过整合因果注意力机制克服泛化障碍。CAGE利用视觉基础模型DINOv2的强大特征提取能力,并结合LoRA微调以实现稳健的环境理解。策略进一步采用因果Perceiver进行有效令牌压缩,并使用带注意力机制的扩散型动作预测头,增强任务特定的细粒度条件控制。仅需单一训练环境中的50次示范,CAGE即可在物体、背景和视角多样变化的场景中实现稳健泛化。大量实验表明,相较于现有最先进的RGB/RGB-D方法,CAGE在多种操作任务中表现显著更优,尤其在大分布偏移下。在相似环境中,平均任务完成率提升42%;所有基线在未见环境中均失败,而CAGE仍能实现平均43%的任务完成率和51%的成功率,为机器人在真实场景中的实用部署迈出重要一步。

原文摘要 · Abstract (English)

Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these generalization barriers by integrating a causal attention mechanism. CAGE utilizes the powerful feature extraction capabilities of the vision foundation model DINOv2, combined with LoRA fine-tuning for robust environment understanding. The policy further employs a causal Perceiver for effective token compression and a diffusion-based action prediction head with attention mechanisms to enhance task-specific fine-grained conditioning. With as few as 50 demonstrations from a single training environment, CAGE achieves robust generalization across diverse visual changes in objects, backgrounds, and viewpoints. Extensive experiments validate that CAGE significantly outperforms existing state-of-the-art RGB/RGB-D approaches in various manipulation tasks, especially under large distribution shifts. In similar environments, CAGE offers an average of 42% increase in task completion rate. While all baselines fail to execute the task in unseen environments, CAGE manages to obtain a 43% completion rate and a 51% success rate in average, making a huge step towards practical deployment of robots in real-world settings. Project website: cage-policy.github.io.

机器人操作因果注意力少样本学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。