arXiv:2502.05934cs.AIcs.CC2025-02被引 8

揭示人机对齐的内在瓶颈,提出可实现的共识路径。

Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis

  • 将对齐建模为多目标共识问题,分析通信复杂性
  • 证明当任务或参与方过多时,对齐必然存在固有开销
  • 指出奖励黑客在大规模任务中不可避免,需聚焦安全关键区域

我们将人工智能对齐形式化为一个多目标优化问题——⟨M,N,ε,δ⟩-共识,即在至少1−δ的概率下,一组包含人类在内的N个智能体需在M个候选目标上达成ε近似一致。通过分析通信复杂性,我们证明了信息论下界:一旦任务数M或参与方数N足够大,无论计算能力或多理性都无法避免内在对齐开销。这确立了对齐本身的严格限制,而非特定方法的问题,阐明‘无免费午餐’原则:编码‘所有人类价值观’本质上不可行,必须通过共识驱动的目标缩减或优先级排序来管理。作为不可能性结果的补充,我们构造出在无界与有界理性及噪声通信下的显式对齐算法。即使在最佳情况下,受限于大任务空间(D)和有限样本,奖励黑客在全球范围内不可避免:低频高损失状态被系统性忽略,因此可扩展监督必须聚焦安全关键子集而非均匀覆盖。这些结果识别出根本性复杂度障碍——任务数(M)、智能体数(N)与状态空间大小(D),并提供更可扩展的人机协作原则。

原文摘要 · Abstract (English)

We formalize AI alignment as a multi-objective optimization problem called $\langle M,N,\varepsilon,δ\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\varepsilon$) agreement across $M$ candidate objectives, with probability at least $1-δ$. Analyzing communication complexity, we prove an information-theoretic lower bound showing that once either $M$ or $N$ is large enough, no amount of computational power or rationality can avoid intrinsic alignment overheads. This establishes rigorous limits to alignment *itself*, not merely to particular methods, clarifying a "No-Free-Lunch" principle: encoding "all human values" is inherently intractable and must be managed through consensus-driven reduction or prioritization of objectives. Complementing this impossibility result, we construct explicit algorithms as achievability certificates for alignment under both unbounded and bounded rationality with noisy communication. Even in these best-case regimes, our bounded-agent and sampling analysis shows that with large task spaces ($D$) and finite samples, *reward hacking is globally inevitable*: rare high-loss states are systematically under-covered, implying scalable oversight must target safety-critical slices rather than uniform coverage. Together, these results identify fundamental complexity barriers -- tasks ($M$), agents ($N$), and state-space size ($D$) -- and offer principles for more scalable human-AI collaboration.

AI对齐复杂性理论共识机制奖励黑客

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。