让机器人理解人类的思考,比单纯问偏好更高效地对齐意图。
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

- 用双层认知模型:人教机器,机器揣摩人对自身的认知。
- 有认知模型的教师比盲目提问效率高30%以上(模拟结果)。
- 适合需要精准对齐的人机协作场景,如医疗或自动驾驶。
当奖励函数无法直接指定时,比较反馈已成为对齐机器人行为与人类意图的标准方法。传统偏好学习将人类视为被动回答者,但此举放弃了人类掌握目标这一核心优势。知晓目标的教师可比学习者主动采样更高效地构建训练数据,且随着奖励特征维度增加,该优势愈发明显。然而,要发挥此优势,需准确掌握学习者当前认知状态。因此,本文将偏好学习重构为一人一机协同问题,耦合两个行为模型:教师维护对学习者的认知模型以设计有效课程;学习者则通过二阶心智理论(ToM-2)模型,生成结构化偏好约束(即‘理解陈述’),使教师的认知模型与自身实际状态保持同步。模拟结果显示,具备认知模型的教师性能优于学习者主导选择;当教师认知出现方向性偏差时,理解陈述能有效修复;且在教师误差集中于特定方向时,二阶心智陈述优于平均信念陈述。
原文摘要 · Abstract (English)
Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher's defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward's feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher's model, emitting structured preference constraints (understanding statements) that keep the teacher's model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher's error about the learner is concentrated in a particular direction rather than spread evenly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。