破解机器学习中的目标扭曲难题,揭示代理指标与真实目标的耦合机制
The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation
- 摆脱独立性假设,研究代理指标与目标间的耦合对优化的影响
- 发现重尾差异下过度优化速率与差异程度成反比
- 适用于各类学习范式,为模型设计提供理论依据
Goodhart定律是政策制定中的著名格言,即当一个度量成为目标时,它便不再是好的度量。随着机器学习模型及其优化能力的发展,越来越多的实证证据支持该定律的有效性,但尚未得到形式化。近期已有少数工作尝试形式化该定律,或对其变体进行分类,或研究优化代理指标如何影响预期目标的优化。本文消除先前工作中依赖独立性的简化假设,以及大多数工作所依赖的学习范式假设,研究代理指标与预期目标之间的耦合对Goodhart定律的影响。结果表明,在目标和偏差均为轻尾分布的情况下,依赖关系不会改变Goodhart效应的本质;然而在目标为轻尾、偏差为重尾的情况下,我们给出一个例子,其中过度优化的发生速率与偏差的重尾程度成反比。
原文摘要 · Abstract (English)
Goodhart's law is a famous adage in policy-making that states that ``When a measure becomes a target, it ceases to be a good measure''. As machine learning models and the optimisation capacity to train them grow, growing empirical evidence reinforced the belief in the validity of this law without however being formalised. Recently, a few attempts were made to formalise Goodhart's law, either by categorising variants of it, or by looking at how optimising a proxy metric affects the optimisation of an intended goal. In this work, we alleviate the simplifying independence assumption, made in previous works, and the assumption on the learning paradigm made in most of them, to study the effect of the coupling between the proxy metric and the intended goal on Goodhart's law. Our results show that in the case of light tailed goal and light tailed discrepancy, dependence does not change the nature of Goodhart's effect. However, in the light tailed goal and heavy tailed discrepancy case, we exhibit an example where over-optimisation occurs at a rate inversely proportional to the heavy tailedness of the discrepancy between the goal and the metric. %
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。