arXiv:2603.05574cs.ROcs.AI2026-03中稿 · publication at Eur…

用人类指令优化机器人抓取策略,提升泛化与效率。

PRISM: Personalized Refinement of Imitation Skills for Manipulation via Human Instructions

  • 基于人类指令迭代生成奖励,融合模仿与强化学习
  • 在模拟环境中提升任务鲁棒性,降低计算开销
  • 支持中间反馈修正,适合需快速适配新场景的机器人

本文提出PRISM:一种基于人类指令的机器人操作模仿策略精细化方法。该方法将模仿学习(IL)与强化学习(RL)无缝整合,使从用户引导示范中生成的通用任务模仿策略,可通过强化学习进一步优化,产生未见过的精细行为。优化过程遵循Eureka范式,从初始自然语言任务描述迭代生成奖励函数。该方法在此基础上,通过引入人类对中间轨迹的反馈修正,实现对通用任务策略在新目标配置下的自适应调整,并加入约束条件,提升策略复用性与数据效率。在模拟抓取-放置任务中的实验表明,相较于无人类反馈的策略,该方法显著提升了部署鲁棒性并减少了计算负担。

原文摘要 · Abstract (English)

This paper presents PRISM: an instruction-conditioned refinement method for imitation policies in robotic manipulation. This approach bridges Imitation Learning (IL) and Reinforcement Learning (RL) frameworks into a seamless pipeline, such that an imitation policy on a broad generic task, generated from a set of user-guided demonstrations, can be refined through reinforcement to generate new unseen fine-grain behaviours. The refinement process follows the Eureka paradigm, where reward functions for RL are iteratively generated from an initial natural-language task description. Presented approach, builds on top of this mechanism to adapt a refined IL policy of a generic task to new goal configurations and the introduction of constraints by adding also human feedback correction on intermediate rollouts, enabling policy reusability and therefore data efficiency. Results for a pick-and-place task in a simulated scenario show that proposed method outperforms policies without human feedback, improving robustness on deployment and reducing computational burden.

机器人操作模仿学习人类反馈策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。