arXiv:2409.15528cs.ROcs.LG2024-09ICRA被引 10

用扩散模型生成多样击打动作,兼顾物理约束与高效学习。

Learning Diverse Robot Striking Motions with Diffusion Models and Kinematically Constrained Gradient Guidance

  • 通过运动学梯度引导,让扩散模型生成符合机器人结构的动作
  • 仿真冰球任务块球率提升25.4%,真实乒乓球成功率达17.3%提升
  • 适合需要快速、多样、受约束动作的机器人学习场景

机器人学习在多样化任务中展现出潜力,但普遍存在样本效率低、难以处理行为多样的数据集,且难以融入物理约束等问题,这对高动态任务如乒乓球尤为重要。现有示范学习方法虽提升了样本效率并支持多样数据,却很少在高动态任务上评估;强化学习则需依赖高保真仿真器。为此,我们提出一种离线、约束引导且能表达多样敏捷行为的扩散建模新方法。核心是运动学约束梯度引导(KCGG)技术,通过机器人正向运动学和扩散模型联合计算梯度,指导采样过程。该方法在不违反约束的前提下,保持生成轨迹与训练数据分布一致。我们在两个挑战性场景中验证了其有效性:模拟冰球与真实乒乓球。在模拟冰球任务中,块球率提升25.4%;在乒乓球任务中,成功率相比模仿学习基线提高17.3%。

原文摘要 · Abstract (English)

Advances in robot learning have enabled robots to generate skills for a variety of tasks. Yet, robot learning is typically sample inefficient, struggles to learn from data sources exhibiting varied behaviors, and does not naturally incorporate constraints. These properties are critical for fast, agile tasks such as playing table tennis. Modern techniques for learning from demonstration improve sample efficiency and scale to diverse data, but are rarely evaluated on agile tasks. In the case of reinforcement learning, achieving good performance requires training on high-fidelity simulators. To overcome these limitations, we develop a novel diffusion modeling approach that is offline, constraint-guided, and expressive of diverse agile behaviors. The key to our approach is a kinematic constraint gradient guidance (KCGG) technique that computes gradients through both the forward kinematics of the robot arm and the diffusion model to direct the sampling process. KCGG minimizes the cost of violating constraints while simultaneously keeping the sampled trajectory in-distribution of the training data. We demonstrate the effectiveness of our approach for time-critical robotic tasks by evaluating KCGG in two challenging domains: simulated air hockey and real table tennis. In simulated air hockey, we achieved a 25.4% increase in block rate, while in table tennis, we saw a 17.3% increase in success rate compared to imitation learning baselines.

机器人学习扩散模型动作生成约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。