arXiv:2502.21262cs.AIcs.LG2025-02

建模人类对AI行为的信念,提升难以监督场景下的价值学习可靠性。

Modeling Human Beliefs about AI Behavior for Scalable Oversight

  • 用数学模型刻画人类对AI行为的认知偏差
  • 发现即使人类误解AI,仍可通过信念覆盖机制实现正确价值学习
  • 适合研究可扩展智能监管与人机对齐的学者

随着AI系统能力超越人类,可扩展的监督变得至关重要:如何监管超出人类能力的AI?一个关键挑战是,人类评估者在复杂任务中可能形成对AI行为的错误认知,导致反馈不可靠,价值推断失效。为此,我们提出建模评估者的信念以更可靠地解释其反馈。通过形式化人类信念模型,分析其在价值学习中的理论作用,并界定残余模糊性的条件。为降低对精确信念模型的依赖,我们引入‘信念模型覆盖’作为松弛策略。这启发了初步方案:利用适配后的基础模型内部表示来模拟人类评估者的信念。这些表示可用于从人类反馈中学习正确价值,即便评估者误解了AI行为。研究表明,建模人类信念能改善价值学习,并为可扩展监督提供了可行的研究路径。

原文摘要 · Abstract (English)

As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex tasks, leading to unreliable feedback and poor value inference. To address this, we propose modeling evaluators' beliefs to interpret their feedback more reliably. We formalize human belief models, analyze their theoretical role in value learning, and characterize when ambiguity remains. To reduce reliance on precise belief models, we introduce "belief model covering" as a relaxation. This motivates our preliminary proposal to use the internal representations of adapted foundation models to mimic human evaluators' beliefs. These representations could be used to learn correct values from human feedback even when evaluators misunderstand the AI's behavior. Our work suggests that modeling human beliefs can improve value learning and outlines practical research directions for implementing this approach to scalable oversight.

价值学习人机对齐可扩展监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。