arXiv:2602.06035cs.CVcs.GR2026-02被引 9

让机器人像人一样自然地操控物体,且能适应新环境和新物品。

InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

  • 通过模仿学习和强化学习,构建统一的生成式控制框架。
  • 在10万次交互数据上训练,可泛化到未见物体和初始状态。
  • 适合需要灵活操作的机器人任务,如人机协作或真实部署。

人类很少以显式的全身动作规划与物体互动。高层意图(如可操作性)定义目标,而平衡、接触和操纵等协调行为可由底层物理与运动先验自然产生。规模化此类先验是实现类人机器人在多样情境下组合与泛化运动-操控技能的关键,同时保持全身动作的物理一致性。为此,本文提出InterPrior,一个通过大规模模仿预训练和强化学习微调的可扩展生成控制框架。该框架首先将全参考模仿专家压缩为一种多功能、目标条件化的变分策略,能够从多模态观测和高层意图中重建运动。尽管该压缩策略能复现训练行为,但在大规模人-物交互配置空间中泛化能力不足。为此,我们引入物理扰动的数据增强,并进行强化学习微调,提升对未见目标和初始状态的适应能力。上述步骤共同将重建的潜在技能整合为有效流形,形成可超越训练数据泛化的运动先验,例如能自然融入与未见物体的交互。我们进一步验证其在用户交互控制中的有效性,并展示了其在真实机器人部署中的潜力。

原文摘要 · Abstract (English)

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors is key to enabling humanoids to compose and generalize loco-manipulation skills across diverse contexts while maintaining physically coherent whole-body coordination. To this end, we introduce InterPrior, a scalable framework that learns a unified generative controller through large-scale imitation pretraining and post-training by reinforcement learning. InterPrior first distills a full-reference imitation expert into a versatile, goal-conditioned variational policy that reconstructs motion from multimodal observations and high-level intent. While the distilled policy reconstructs training behaviors, it does not generalize reliably due to the vast configuration space of large-scale human-object interactions. To address this, we apply data augmentation with physical perturbations, and then perform reinforcement learning finetuning to improve competence on unseen goals and initializations. Together, these steps consolidate the reconstructed latent skills into a valid manifold, yielding a motion prior that generalizes beyond the training data, e.g., it can incorporate new behaviors such as interactions with unseen objects. We further demonstrate its effectiveness for user-interactive control and its potential for real robot deployment.

生成控制人机交互强化学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。