arXiv:2607.00155cs.AIcs.GT2026-07

研究人机协作中双向隐私信息下的监督机制,揭示信任盲区与可避免伤害。

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

  • 构建双向信息不对称的上下文老虎机监督游戏,模拟真实人机协作场景
  • 发现贪婪人类因过度信任先验而拒绝监督,导致可避免伤害区域存在
  • 提出通过延迟响应信号实现动态学习,缓解非可信监督沟通问题

我们研究在运行时由人类进行监督的AI代理情境,其中双方都拥有私有信息:人类私密地知晓其奖励函数,而AI私密地知晓所提行动的质量。这种双向信息不对称自然出现在自主机器人或软件代理评估人类无法直接观察的情境中。基于合作逆强化学习(CIRL)和监督博弈,我们引入一种具有双向信息不对称的上下文老虎机团队博弈,并采用‘执行/请求/信任/监督’接口。该老虎机结构消除了物理状态转移,从而获得在完整部分可观马尔可夫决策过程(POMDP)设定下仅能假设的一次性刻画,尽管共同信念仍是跨轮次动态控制的状态。我们给出了两个一次性刻画:团队最优解和行为上自然的短视规则,两者之间的差距构成可避免伤害的区域:在此区域内,AI私下知道提议动作有害,关闭系统有益,但短视的人类因信任其先验而拒绝监督。我们证明这一差距是不可信监督通信的代价,并对通过被动学习和一周期滞后的监督响应实现的动态化解进行了初步分析。

原文摘要 · Abstract (English)

We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of asymmetry that arises naturally when an autonomous robot or software agent has inspected a situation its human supervisor cannot directly assess. Building on Cooperative Inverse Reinforcement Learning (CIRL) and the Oversight Game, we introduce a contextual-bandit team game with two-sided asymmetric information and a play/ask/trust/oversee interface. The bandit structure removes physical state transitions and thereby yields exact one-shot characterizations that would remain conjectural in the full POMDP setting, though the common belief remains a dynamically controlled state across rounds. We give two one-shot characterizations, a team optimum and a behaviorally natural myopic rule, whose gap is a slab of avoidable harm: a region in which the AI privately knows the proposed action is harmful and shutdown would help, yet a myopic human, trusting her prior, declines to oversee. We show this gap is the price of non-credible oversight communication, and give a partial analysis of how it resolves dynamically over repeated rounds through passive learning and active signaling with a one-period-lagged oversight response.

人机协作信息不对称监督博弈逆强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。