arXiv:2507.05695cs.RO2025-07中稿 · ICRA被引 4

用几何代数让机器人学得更快更准

Hybrid Diffusion Policies with Projective Geometric Algebra for Efficient Robot Manipulation Learning

  • 把几何变换知识直接嵌入网络结构,避免重复学习
  • 训练速度比传统方法快得多,真实场景表现优异
  • 适合需要高效空间推理的机器人控制任务

扩散策略是机器人学习的强大范式,但训练效率常不高。主要原因是网络需为每个新任务从头学习基本空间概念(如平移、旋转)。为此,我们提出将投影几何代数(PGA)的几何归纳偏置直接嵌入网络架构。PGA提供统一的代数框架,用于表示几何原语和变换,使神经网络更有效地推理空间结构。本文提出hPGA-DP,一种新型混合扩散策略:采用基于PGA的Transformer(P-GATr)作为状态编码器和动作解码器,同时使用成熟的U-Net或Transformer模块完成核心去噪过程。在仿真与真实环境中的大量实验及消融研究显示,hPGA-DP显著提升任务性能与训练效率。特别地,该混合方法相比标准扩散策略及仅依赖P-GATr的架构,收敛速度明显更快。

原文摘要 · Abstract (English)

Diffusion policies are a powerful paradigm for robot learning, but their training is often inefficient. A key reason is that networks must relearn fundamental spatial concepts, such as translations and rotations, from scratch for every new task. To alleviate this redundancy, we propose embedding geometric inductive biases directly into the network architecture using Projective Geometric Algebra (PGA). PGA provides a unified algebraic framework for representing geometric primitives and transformations, allowing neural networks to reason about spatial structure more effectively. In this paper, we introduce hPGA-DP, a novel hybrid diffusion policy that capitalizes on these benefits. Our architecture leverages the Projective Geometric Algebra Transformer (P-GATr) as a state encoder and action decoder, while employing established U-Net or Transformer-based modules for the core denoising process. Through extensive experiments and ablation studies in both simulated and real-world environments, we demonstrate that hPGA-DP significantly improves task performance and training efficiency. Notably, our hybrid approach achieves substantially faster convergence compared to both standard diffusion policies and architectures that rely solely on P-GATr. The project website is available at: https://apollo-lab-yale.github.io/26-ICRA-hPGA-website/.

机器人学习扩散模型几何代数高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。