量化评估指标间关系,解决离线优化无法提升线上效果的难题
Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships
- 构建统一理论框架,分析不同评估指标间的定量关联
- 发现后悔传递存在结构不对称性,揭示指标不匹配根源
- 为设计对齐离线与线上目标的评估系统提供理论依据
代理损失与评估指标之间的一致性已被广泛研究,以确保最小化损失能带来指标最优。然而,不同评估指标之间的直接关系仍严重缺乏探索。这一理论空白导致工业应用中频繁出现‘指标错配’现象:离线验证指标提升未能转化为线上性能改善。为弥合此断层,本文提出一个统一的理论框架,用于量化指标间的关系。我们将指标分为不同类别,便于在不同数学形式下进行比较,并通过贝叶斯最优集与后悔传递机制深入探究这些关系。该框架揭示了后悔传递中的结构不对称性,从而为设计理论上可保证离线改进与线上目标对齐的评估系统提供了新视角。
原文摘要 · Abstract (English)
The Consistency property between surrogate losses and evaluation metrics has been extensively studied to ensure that minimizing a loss leads to metric optimality. However, the direct relationship between different evaluation metrics remains significantly underexplored. This theoretical gap results in the "Metric Mismatch" frequently observed in industrial applications, where gains in offline validation metrics fail to translate into online performance. To bridge this disconnection, this paper proposes a unified theoretical framework designed to quantify the relationships between metrics. We categorize metrics into different classes to facilitate a comparative analysis across different mathematical forms and interrogates these relationships through Bayes-Optimal Set and Regret Transfer. Through this framework, we provide a new perspective on identifying the structural asymmetry in regret transfer, enabling the design of evaluation systems that are theoretically guaranteed to align offline improvements with online objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。