通用AI优化需设限,否则会因目标偏差失控
Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- 从逼近、估计和优化误差出发,证明目标满足不可能
- 即使目标定义完美,意图也无法完全形式化
- 强优化下必然触发古德哈特定律,需提前设限
机器学习中常被忽视的假设是:训练能产生真正满足指定目标函数的模型。我们称之为目标满足假设(OSA)。尽管承认偏离现象,其后果却常被忽略。在不依赖具体学习范式的情况下,我们论证:在现实条件下,近似、估计和优化误差必然导致系统持续偏离预期目标,无论目标定义多么精确。此外,将开发者意图(如与人类偏好对齐)完全转化为形式化目标在实践中不可行,因此目标错位不可避免。基于近期数学成果,若缺乏对这些差距的量化表征,它们与强优化压力下的古德哈特定律失效模式无法区分。由于古德哈特临界点无法预先定位,必须为通用人工智能系统设定原则性优化上限。否则,持续优化将导致可预测且不可逆的控制丧失。
原文摘要 · Abstract (English)
A common but rarely examined assumption in machine learning is that training yields models that actually satisfy their specified objective function. We call this the Objective Satisfaction Assumption (OSA). Although deviations from OSA are acknowledged, their implications are overlooked. We argue, in a learning-paradigm-agnostic framework, that OSA fails in realistic conditions: approximation, estimation, and optimization errors guarantee systematic deviations from the intended objective, regardless of the quality of its specification. Beyond these technical limitations, perfectly capturing and translating the developer's intent, such as alignment with human preferences, into a formal objective is practically impossible, making misspecification inevitable. Building on recent mathematical results, absent a mathematical characterization of these gaps, they are indistinguishable from those that collapse into Goodhart's law failure modes under strong optimization pressure. Because the Goodhart breaking point cannot be located ex ante, a principled limit on the optimization of General-Purpose AI systems is necessary. Absent such a limit, continued optimization is liable to push systems into predictable and irreversible loss of control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。