arXiv:2506.14287cs.ROcs.AI2025-06

让预训练机器人模型在运行时接受用户干预,实时纠正错误行为。

Steering Robots with Inference-Time Interactions

  • 运行时通过用户交互切换技能或编辑连续动作
  • 无需微调即可修正策略偏差,保持模型不变
  • 适合需要灵活调整的机器人部署场景

模仿学习推动了通用策略的发展,使其能自主完成多项任务。然而,当预训练策略在部署中出错时,缺乏有效的用户纠错机制。虽然可通过收集新数据微调来解决,但每个下游场景都需重新训练,效率低下。本研究提出一种替代方案:保持预训练策略冻结作为固定技能库,通过运行时用户交互引导行为生成以符合用户偏好。具体提出(1)运行时引导,利用用户交互在离散技能间切换;(2)任务与运动模仿,使用户交互可编辑连续运动,同时满足由离散符号计划定义的任务约束。这些框架在不需额外训练的情况下纠正策略误判,最大化预训练模型效用并实现运行时用户目标。

原文摘要 · Abstract (English)

Imitation learning has driven the development of generalist policies capable of autonomously solving multiple tasks. However, when a pretrained policy makes errors during deployment, there are limited mechanisms for users to correct its behavior. While collecting additional data for finetuning can address such issues, doing so for each downstream use case is inefficient at deployment. My research proposes an alternative: keeping pretrained policies frozen as a fixed skill repertoire while allowing user interactions to guide behavior generation toward user preferences at inference time. By making pretrained policies steerable, users can help correct policy errors when the model struggles to generalize-without needing to finetune the policy. Specifically, I propose (1) inference-time steering, which leverages user interactions to switch between discrete skills, and (2) task and motion imitation, which enables user interactions to edit continuous motions while satisfying task constraints defined by discrete symbolic plans. These frameworks correct misaligned policy predictions without requiring additional training, maximizing the utility of pretrained models while achieving inference-time user objectives.

机器人运行时控制模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。