arXiv:2605.22986cs.ROcs.AI2026-05

让机器人主动问问题,修复演示数据中的奖励遗漏。

Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations

论文配图:Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
图 1 · 摘自论文原文
  • 通过分析演示数据波动性,识别未充分演示的特征。
  • 用自然语言解释不确定点,精准请求纠正演示,提升奖励学习效果。
  • 适合需要高可靠性的人机协作场景,如真实机器人操作。

从示范中学习奖励函数通常假设示范覆盖了所有重要特征,但实践中示范常因认知负荷或物理困难而忽略某些方面,导致关键特征未被充分说明,从而在部署时引发行为偏差。本文提出一种框架,可检测这些未充分指定的特征,并主动发起针对性的纠正示范请求。核心思路是:若某特征在多个示范中变化小,说明其已被充分表达;反之,变化大则可能未被充分演示。利用这一统计信号,机器人能推断出不确定性所在,并以自然语言解释其困惑,向人类请求聚焦于这些缺陷的示范。我们在模拟桌面操作环境和真实Franka机械臂用户研究中验证该方法,结果表明,基于解释的定向查询显著优于随机提问与被动数据收集,有效减少了由不完美示范带来的奖励歧义。

原文摘要 · Abstract (English)

Learning reward functions from demonstrations assumes that demonstrations provide adequate supervision over all features -- or task-relevant aspects of behavior. In practice, demonstrations are often imperfect: humans may under-emphasize certain features due to cognitive load or physical difficulty, or the training regime may fail to sufficiently cover all relevant situations. In either case, important features may be underspecified, leading to ambiguity in the learned reward function and misaligned behavior at deployment. We propose a framework that detects such underspecified features and actively solicits targeted corrective demonstrations. Our key insight is that demonstrations implicitly reveal which features are well specified: features that are consistently optimized show little variation across demonstrations, while features that are underspecified vary widely. We leverage this statistical signal to infer which features may have been insufficiently demonstrated. The robot then explains which features it is uncertain about in natural language and queries for demonstrations that explicitly address the identified gaps. We evaluate our approach in a simulated tabletop manipulation domain and in a user study with a real Franka robot. Targeted, explanation-guided queries significantly improve reward recovery compared to random querying and passive data collection, reducing ambiguity that would otherwise persist in learning from imperfect demonstrations.

强化学习人机交互机器人奖励学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。