arXiv:2512.15584cs.AIcs.GT2025-12

提出决策理论框架,判断何时可放心把决策交给不完美对齐的AI。

A Decision-Theoretic Approach for Managing Misalignment

  • 用决策论建模价值错位、认知准确性和行动范围的权衡
  • 发现特定情境下即使严重错位也能理性委托,因精度或能力提升收益更大
  • 适合关注AI治理与风险控制的研究者和实践者

在不确定性下,何时应将决策权交给AI系统?尽管价值对齐研究已发展出塑造AI价值观的方法,但对如何判断不完美对齐是否足以授权仍关注不足。本文提出一种形式化决策理论框架,精确分析代理者价值(错)对齐度、认知准确性及可行动范围之间的权衡,并考虑委托方对此三者的不确定性。分析揭示两种委托情景的本质差异:第一,普遍性委托(信任代理处理任何问题)要求近乎完美的价值对齐与完全认知信任,实践中极少满足;第二,我们证明,在特定情境下,即使存在显著价值错位,若代理具备更高精度或更广行动范围,仍可能带来更优决策结果,使委托在期望上合理。为此,我们设计了一种新颖的评分机制,用于事前量化决策合理性。最终工作提供了一种原则性方法,判断某个AI在特定情境中是否对齐到足够程度,推动从追求完美对齐转向在不确定性下管理委托的风险与收益。

原文摘要 · Abstract (English)

When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good enough to justify delegation. We argue that rational delegation requires balancing an agent's value (mis)alignment with its epistemic accuracy and its reach (the acts it has available). This paper introduces a formal, decision-theoretic framework to analyze this tradeoff precisely accounting for a principal's uncertainty about these factors. Our analysis reveals a sharp distinction between two delegation scenarios. First, universal delegation (trusting an agent with any problem) demands near-perfect value alignment and total epistemic trust, conditions rarely met in practice. Second, we show that context-specific delegation can be optimal even with significant misalignment. An agent's superior accuracy or expanded reach may grant access to better overall decision problems, making delegation rational in expectation. We develop a novel scoring framework to quantify this ex ante decision. Ultimately, our work provides a principled method for determining when an AI is aligned enough for a given context, shifting the focus from achieving perfect alignment to managing the risks and rewards of delegation under uncertainty.

AI治理决策理论对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。