arXiv:2512.17180cs.ROcs.AI2025-12中稿 · ACRA 2025

AI agents更信任低奖励教师,而非高奖励者。

Conservative Bias in Multi-Teacher Learning: Why Agents Prefer Low-Reward Advisors

  • 实验发现93.16%的智能体偏好低奖励教师,因追求行为一致性。
  • 当教师可用性≥0.6且准确率≥0.6时系统才稳定,否则崩溃。
  • 在概念漂移下比基线Q-learning提升159%,适合安全关键场景。

交互式强化学习(IRL)使自主代理和机器人能从人类教师处学习复杂行为,但教师选择机制尚不明确。本研究揭示:在多个具有不同奖励结构的教师中,学习智能体93.16%的概率选择低奖励教师,而非奖励高出20倍的教师。基于1,250次导航任务实验,发现:(1)保守偏好主导教师选择,智能体优先选择最低奖励教师以保障行为一致性;(2)当教师可用性ρ≥0.6且准确率ω≥0.6时系统才能正常运行,低于此阈值将导致灾难性失败;(3)在概念漂移条件下,该框架相较基线Q-learning性能提升159%。这些结果挑战了强化学习中关于最优教学的基本假设,提示人类对安全与一致性的偏好可能与智能体选择行为相符,为安全关键型机器人训练提供新思路。

原文摘要 · Abstract (English)

Interactive reinforcement learning (IRL) has shown promise in enabling autonomous agents and robots to learn complex behaviours from human teachers, yet the dynamics of teacher selection remain poorly understood. This paper reveals an unexpected phenomenon in IRL: when given a choice between teachers with different reward structures, learning agents overwhelmingly prefer conservative, low-reward teachers (93.16% selection rate) over those offering 20x higher rewards. Through 1,250 experimental runs in navigation tasks with multiple expert teachers, we discovered: (1) Conservative bias dominates teacher selection: agents systematically choose the lowest-reward teacher, prioritising consistency over optimality; (2) Critical performance thresholds exist at teacher availability rho >= 0.6 and accuracy omega >= 0.6, below which the framework fails catastrophically; (3) The framework achieves 159% improvement over baseline Q-learning under concept drift. These findings challenge fundamental assumptions about optimal teaching in RL and suggest potential implications for human-robot collaboration, where human preferences for safety and consistency may align with the observed agent selection behaviour, potentially informing training paradigms for safety-critical robotic applications.

强化学习教师选择安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。