解决模型筛选后评估偏差问题,实现可靠性能-可靠性权衡分析
Post-Selection Distributional Model Evaluation
- 基于e-values构建筛选后分布评估框架,避免选择偏差
- 在真实场景中验证了多层级可靠性下模型性能的可信比较
- 适合需探索性能与可靠性平衡的研究者和工程团队
传统模型评估方法通常用于验证模型是否达到预设的关键绩效指标(KPI)水平。但在许多应用中,目标KPI水平事先未知,用户更希望在测试时全面分析模型在性能与可靠性之间的权衡。这一任务需要对测试阶段的KPI分布进行可靠估计,但通常需使用相同数据既预选模型又估计分布,导致潜在的后选择偏差。本文提出后选择分布模型评估(PS-DME),一种通用框架,可在任意数据依赖的模型预选后进行统计有效的分布评估。基于e-values,PS-DME控制后选择错误覆盖率(FCR),并建立了其相比样本分割基线方法更高效的具体条件。在合成数据、大语言模型文本到SQL解码及电信网络性能评估中的实验表明,PS-DME能可靠地跨多个可靠性水平比较候选配置,支持对性能-可靠性权衡的统计可信赖探索。
原文摘要 · Abstract (English)
Formal model evaluation methods typically certify that a model satisfies a prescribed target key performance indicator (KPI) level. However, in many applications, the relevant target KPI level may not be known a priori, and the user may instead wish to compare candidate models by analyzing the full trade-offs between performance and reliability achievable at test time by the models. This task, requiring the reliable estimate of the test-time KPI distributions, is made more complicated by the fact that the same data must often be used both to pre-select a subset of candidate models and to estimate their KPI distributions, causing a potential post-selection bias. In this work, we introduce post-selection distributional model evaluation (PS-DME), a general framework for statistically valid distributional model assessment after arbitrary data-dependent model pre-selection. Building on e-values, PS-DME controls post-selection false coverage rate (FCR) for the distributional KPI estimates and we establish explicit conditions under which it is provably more sample efficient than a baseline method based on sample splitting. Experiments on synthetic data, text-to-SQL decoding with large language models, and telecom network performance evaluation demonstrate that PS-DME enables reliable comparison of candidate configurations across a range of reliability levels, supporting the statistically reliable exploration of performance--reliability trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。