arXiv:2607.28881cs.AI2026-07

研究过度优化如何破坏人类价值,警告现有对齐方法的脆弱性。

Fragility of Value under Imperfect Alignment

  • 构建理想对齐训练模型,分析代理目标与真实价值的偏差
  • 发现当代理条件误差超过阈值时,优化会引发η级灾难性后果
  • 支持限制优化压力的设计,如量化器,而非依赖训练阶段对齐

随着人工智能系统承担越来越多的责任,确保其与人类价值观对齐变得愈发重要。一个常见的担忧是:人类价值观具有脆弱性——即过度优化某个不完美的价值代理,可能导致灾难性结果。本文提出一种理想化的对齐问题模型,其中智能体在优化世界前经过对齐训练,保证其价值函数满足某种代理条件。主要结果识别了在人类价值函数和若干代理条件准确度满足特定条件下,一个具备η-灾难性价值函数(即在优化能力趋于无穷时,人类价值期望低于η)的智能体仍可能被部署。研究揭示了过度优化的风险,强调应设计限制优化压力的机制(如量化器),而非仅依赖部署前的训练。

原文摘要 · Abstract (English)

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $η$-catastrophic value function, one that is guaranteed to take the expectation of human value below $η$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

AI对齐价值脆弱性优化压力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。