arXiv:2603.17300cs.RO2026-03被引 3

提升机器人多任务指令响应能力,让其能随时听从新指令切换动作。

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

  • 通过分析任务轨迹重叠度,量化并识别机器人缺乏指令响应能力的问题。
  • 在模拟环境中使指令响应能力提升11%,真实世界中实现即时指令切换。
  • 适合研究机器人交互、自主决策与强化学习策略优化的开发者参考。

尽管具备强大的多任务预训练能力,现有机器人策略仍普遍存在任务可引导性差的问题。例如,当机器人正走向烤箱时,若收到新指令‘把碗放进水槽’,它可能仍执行‘关烤箱’动作,尽管单独执行两项任务均无问题。本文提出ReSteer框架,用于量化和提升多任务机器人策略的可引导性。通过对当前先进策略的全面评估,发现可引导性差与训练任务轨迹分布重叠度低密切相关,并引入一种基于策略行为的代理指标来衡量该重叠度。基于此,ReSteer包含三个组件:(i) 不需完整回放即可识别低可引导性状态的引导性估计算法;(ii) 从这些状态下合成运动片段的数据生成器;(iii) 利用生成数据进行自优化的策略精炼流程。在LIBERO仿真环境中,经过18,000次回放,可引导性提升11%。真实世界实验表明,提升可引导性对交互式使用至关重要,使用户可在任意时刻下达任一任务指令。本工作希望推动对可引导性量化及大规模机器人策略数据收集策略的进一步研究。

原文摘要 · Abstract (English)

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the sink" when moving towards the oven, executing ``close the oven", even though it can complete both tasks when executed separately. We propose ReSteer, a framework to quantify and improve task steerability in multitask robot policies. We conduct an exhaustive evaluation of state-of-the-art policies, revealing a common lack of steerability. We find that steerability is associated with limited overlap among training task trajectory distributions, and introduce a proxy metric to measure this overlap from policy behavior. Building on this insight, ReSteer improves steerability via three components: (i) a steerability estimator that identifies low-steerability states without full-rollout evaluation, (ii) a steerable data generator that synthesizes motion segments from these states, and (iii) a self-refinement pipeline that improves policy steerability using the generated data. In simulation on LIBERO, ReSteer improves steerability by 11\% over 18k rollouts. In real-world experiments, we show that improved steerability is critical for interactive use, enabling users to instruct robots to perform any task at any time. We hope this work motivates further study on quantifying steerability and data collection strategies for large robot policies.

机器人控制指令引导策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。