研究指标优化如何扭曲真实目标,揭示了衡量标准失效的深层机制。
On Goodhart's law, with an application to value alignment
- 通过分析目标与指标的偏差尾部分布,揭示优化指标可能偏离真实目标
- 发现长期尾部偏差会导致强型好希特定律:过度优化反而损害目标
- 为基于指标的政策与自动化决策系统提供安全性评估新思路
当一个度量成为目标时,它便不再是一个好的度量——这便是著名的‘好希特定律’。本文从形式上研究该定律,证明其关键取决于真实目标与被优化度量之间偏差的尾部分布。具有长尾分布的偏差会强化好希特定律,即优化度量可能导致对目标的反效果。本文提出一个正式框架,通过研究度量优化时目标与度量相关性的渐近行为来评估好希特定律。此外,区分了‘弱型’好希特定律(过度优化指标对目标无益)与‘强型’好希特定律(过度优化指标有害于目标),并证明二者差异由尾部分布决定。本文强调该结果对大规模决策及依赖度量的政策的重要意义,尤其针对算法自动化的场景,提出多项研究方向以更好评估此类政策的安全性。
原文摘要 · Abstract (English)
``When a measure becomes a target, it ceases to be a good measure'', this adage is known as {\it Goodhart's law}. In this paper, we investigate formally this law and prove that it critically depends on the tail distribution of the discrepancy between the true goal and the measure that is optimized. Discrepancies with long-tail distributions favor a Goodhart's law, that is, the optimization of the measure can have a counter-productive effect on the goal. We provide a formal setting to assess Goodhart's law by studying the asymptotic behavior of the correlation between the goal and the measure, as the measure is optimized. Moreover, we introduce a distinction between a {\it weak} Goodhart's law, when over-optimizing the metric is useless for the true goal, and a {\it strong} Goodhart's law, when over-optimizing the metric is harmful for the true goal. A distinction which we prove to depend on the tail distribution. We stress the implications of this result to large-scale decision making and policies that are (and have to be) based on metrics, and propose numerous research directions to better assess the safety of such policies in general, and to the particularly concerning case where these policies are automated with algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。