提出新评估方法,发现传统验证法存在噪声,影响结果可信度。
Beyond the Fold: Quantifying Split-Level Noise and the Case for Leave-One-Dataset-Out AU Evaluation
- 用留一数据集外法替代传统交叉验证,减少随机波动
- 实测平均F1噪声达±0.065,低频AU波动更显著
- 新方法能暴露领域不稳定性,结果更可靠适合对比
面部动作单元(AU)检测的主流评估协议为受试者专属交叉验证,但报告的性能提升常较小。我们发现该验证方法本身引入可测量的随机方差:在BP4D+数据集上,重复3折受试者专属划分产生±0.065的平均F1噪声,低频AU变异更大。操作点指标如F1比阈值无关的AUC波动更剧烈,模型排名会随划分不同而改变。进一步采用跨五数据集的留一数据集外(LODO)协议评估跨数据集鲁棒性,该方法消除划分随机性,揭示了单数据集交叉验证下无法观察到的领域级不稳定性。结果表明,许多报道的交叉验证提升可能仅在协议噪声范围内。留一数据集外验证提供更稳定、可解释的结果。
原文摘要 · Abstract (English)
Subject-exclusive cross-validation is the standard evaluation protocol for facial Action Unit (AU) detection, yet reported improvements are often small. We show that cross-validation itself introduces measurable stochastic variance. On BP4D+, repeated 3-fold subject-exclusive splits produce an empirical noise floor of $\pm 0.065$ in average F1, with substantially larger variation for low-prevalence AUs. Operating-point metrics such as F1 fluctuate more than threshold-independent measures such as AUC, and model ranking can change under different fold assignments. We further evaluate cross-dataset robustness using a Leave-One-Dataset-Out (LODO) protocol across five AU datasets. LODO removes partition randomness and exposes domain-level instability that is not visible under single-dataset cross-validation. Together, these results suggest that gains often reported in cross-fold validation may fall within protocol variance. Leave-one-dataset-out cross-validation yields more stable and interpretable findings
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。