用训练日志调整模型对比,可降低不确定性,但选错日志会适得其反。
Can Training Logs Make Model Comparisons More Precise?

- 用每模型自身训练日志做针对性调整,不混用数据。
- 早期日志调整可显著减少模型对比的不确定性。
- 适合关注模型差异精确性、有训练日志的实验者。
比较随机训练模型需估计性能差异及其不确定性,来自重复运行的训练日志能否提升精度?由于日志在训练中生成而非训练前测量,采用模型专属协变量调整:仅用各模型自身运行的统计量进行调整,原始均值差仍作为报告效应。在涵盖三种架构和三个数据集的视觉研究中,基于早期训练日志的简单调整常能有效降低模型比较的不确定性。主要限制在于协变量选择:广泛搜索日志池寻找最相关统计量,即便事后发现有用,也可能引入更多噪声。因此,训练日志对更精确的模型比较有用,但前提是调整过程避免大规模选择噪声。
原文摘要 · Abstract (English)
Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether training logs from those same runs can make such comparisons more precise. Because training-log covariates are produced during training rather than measured before it, we use arm-specific covariate adjustment: each model is adjusted only with statistics from its own runs, and the raw mean difference remains the reported effect. In a vision study spanning three architectures and three datasets, simple adjustments based on early training logs often reduce uncertainty in model comparisons. The main limitation is covariate selection. Broadly searching the log pool for the most correlated statistic often adds more noise than it removes, even when useful statistics exist in hindsight. Training logs therefore appear useful for more precise model comparisons, but only when the adjustment avoids large selection noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。