arXiv:2605.19557stat.MLcs.LG2026-05

通过密度比损失实现无需重训练的可调延迟决策

Density-Ratio Losses for Post-Hoc Learning to Defer

  • 基于理想分布的密度比估计,构建可后置调整的延迟评分器
  • 在多种数据集上表现优于基线,且对分布变化更鲁棒
  • 适合需要灵活延迟策略的高风险场景应用

我们从理想分布的角度研究后置学习延迟(L2D),通过在模型损失最低条件下对数据分布进行偏差正则化重加权来定义。通过模型与专家理想分布之间的密度比定义延迟行为。利用密度比估计到分类概率估计的归约方法,推导出适用于后置L2D评分器的DR CPE损失。延迟决策通过阈值化评分器实现,可在不重新训练的情况下调节延迟率。对于基于KL的理想分布,我们的延迟规则在原始分布下恢复了Chow规则,并在联合或边际理想分布下分别建立与专家倾斜贝叶斯后验的联系——该后验融合了专家性能。实验表明,该方法在多个基准中具有竞争力,且在不同数据设置下更稳健。更广泛而言,本研究将后置L2D视为理想分布间的密度比学习,连接了Chow式规则、专家比较,并揭示了与异常检测等学习任务的深层关联。

原文摘要 · Abstract (English)

We study post-hoc Learning to Defer (L2D) through the lens of ideal distributions: divergence-regularized reweightings of the data distribution under which a model attains low loss. We define deferral via the density-ratio between a model's and an expert's ideals. Using the reduction from density-ratio estimation to class-probability estimation, we derive the DR CPE losses for post-hoc L2D scorers. Deferral decisions are then made by thresholding the scorer, allowing deferral rates to be adjusted without retraining. For KL-based ideal distributions, our deferral rules recovers Chow's rule under the original distribution and a connection to an expert-tilted Bayes posterior -- which incorporates the expert's performance -- depending on if the ideal distributions are joint or marginal distributions. Experimentally, our approach is competitive compared to common baselines and more robust across dataset settings. More broadly, our results cast post-hoc L2D as density-ratio learning between ideal distributions, bridging Chow-style rules, expert comparison, and elucidating connections to related learning settings including anomaly detection.

学习延迟密度比可调决策专家融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。