将机器人运动结构先验融入神经网络,提升动作学习效果。
Rodrigues Network for Learning Robot Actions
- 用可学习的罗德里格斯算子替代传统网络,引入运动学先验
- 在合成任务中显著优于标准模型,在真实场景中提升模仿学习性能
- 适合需要精准动作建模的机器人控制与三维重建任务
理解与预测关节动作在机器人学习中至关重要。然而,常见架构如MLP和Transformer缺乏反映关节系统内在运动学结构的归纳偏置。为此,我们提出神经罗德里格斯算子,一种可学习的前向运动学推广形式,旨在向神经计算注入运动学感知的归纳偏置。基于该算子,设计了专用于处理动作的罗德里格斯网络(RodriNet)。我们在两个合成任务上评估了网络的表达能力,结果显示其在运动学预测和运动预测任务中均显著优于标准主干网络。此外,我们在两个真实应用场景中验证了其有效性:(i) 在机器人基准上使用扩散策略进行模仿学习;(ii) 单图三维手部重建。结果表明,将结构化运动学先验融入网络架构能显著提升多领域中的动作学习表现。
原文摘要 · Abstract (English)
Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the underlying kinematic structure of articulated systems. To this end, we propose the Neural Rodrigues Operator, a learnable generalization of the classical forward kinematics operation, designed to inject kinematics-aware inductive bias into neural computation. Building on this operator, we design the Rodrigues Network (RodriNet), a novel neural architecture specialized for processing actions. We evaluate the expressivity of our network on two synthetic tasks on kinematic and motion prediction, showing significant improvements compared to standard backbones. We further demonstrate its effectiveness in two realistic applications: (i) imitation learning on robotic benchmarks with the Diffusion Policy, and (ii) single-image 3D hand reconstruction. Our results suggest that integrating structured kinematic priors into the network architecture improves action learning in various domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。