arXiv:2602.18603cs.ROcs.LG2026-02

通过纠正时机提升机器人对任务目标的理解能力。

Enhancing Goal Inference via Correction Timing

  • 利用人类纠正行为的时间点作为反馈信号,捕捉其决策依据。
  • 基于纠正时机与初始方向,可快速推断出最终任务目标。
  • 适用于人机协作场景中需快速理解意图的机器人学习系统。

纠正行为为人类向机器人提供反馈提供了自然方式:一方面在认为机器人将失败时介入,另一方面修改机器人行为以完成任务。每次纠正都传递了应做与不应做的信息,且修正后的行为比原始行为更符合任务目标。现有大多数学习纠正的方法将其视为新示范或偏好比较,却忽略了人类决定干预的关键因素——纠正时机。该时机可能受任务进展、人类预期、动态特性、动作可读性及最优性等多种因素影响。本文研究纠正时机是否能作为有用信号来推断这些任务相关影响。具体探索三个应用:(1) 识别引发人类纠正的运动特征;(2) 仅凭纠正时机与初始方向快速推断最终目标;(3) 学习更精确的任务约束。结果表明,纠正时机显著提升了前两项应用的学习效果。整体上,本工作揭示了纠正时机在机器人学习中的新价值。

原文摘要 · Abstract (English)

Corrections offer a natural modality for people to provide feedback to a robot, by (i) intervening in the robot's behavior when they believe the robot is failing (or will fail) the task objectives and (ii) modifying the robot's behavior to successfully fulfill the task. Each correction offers information on what the robot should and should not do, where the corrected behavior is more aligned with task objectives than the original behavior. Most prior work on learning from corrections involves interpreting a correction as a new demonstration (consisting of the modified robot behavior), or a preference (for the modified trajectory compared to the robot's original behavior). However, this overlooks one essential element of the correction feedback, which is the human's decision to intervene in the robot's behavior in the first place. This decision can be influenced by multiple factors including the robot's task progress, alignment with human expectations, dynamics, motion legibility, and optimality. In this work, we investigate whether the timing of this decision can offer a useful signal for inferring these task-relevant influences. In particular, we investigate three potential applications for this learning signal: (1) identifying features of a robot's motion that may prompt people to correct it, (2) quickly inferring the final goal of a human's correction based on the timing and initial direction of their correction motion, and (3) learning more precise constraints for task objectives. Our results indicate that correction timing results in improved learning for the first two of these applications. Overall, our work provides new insights on the value of correction timing as a signal for robot learning.

人机协作纠正学习任务理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。