提出跨域不确定性量化新方法,提升小样本下选择性预测的可靠性。
Cross-Domain Uncertainty Quantification for Selective Prediction: A Comprehensive Bound Ablation with Transfer-Informed Betting
- 用迁移信息初始化赌博信心序列,增强小样本时的边界紧致性。
- 在MASSIVE数据集上实现94.0%覆盖精度,比传统方法提升27%。
- 适合需要高置信度单预测保障的智能缓存系统等场景。
我们系统评估了九种有限样本界家族在带风险控制的选择性预测中的表现,结合集中不等式(Hoeffding、Empirical Bernstein、Clopper-Pearson、Wasserstein DRO、CVaR)、多重检验校正(并集界、学习后测试固定序列)与基于赌博的信心序列(WSR)。主要理论贡献是迁移感知赌博(TIB),通过源域风险分布预热WSR财富过程,在数据稀缺条件下获得更紧的边界,并具有形式化优势保证。证明TIB财富过程在任意源-目标差异下仍为合法超鞅,当领域匹配时优于标准WSR,且不存在数据无关的预热方式能实现更好收敛。该三者结合——赌博信心序列、LTT单调检验、跨域迁移——在现有文献中尚属首次。我们在四个基准上评估:MASSIVE(n=1,102)、NyayaBench(n=280)、CLINC-150(n=22.5K)、Banking77(n=13K),共18组(alpha, delta)配置。在MASSIVE上,alpha=0.10时,LTT消除了ln(K)并集界惩罚,达到94.0%保证覆盖,远超Hoeffding的73.8%(相对提升27%)。在NyayaBench上,因校准集过小,传统Hoeffding类边界在alpha<0.20时不可行,而TIB在alpha=0.10时实现18.5%覆盖,较LTT+Hoeffding提升5.4倍。对比分层置信预测,发现其生成平均1.67类的预测集,而选择性预测提供单预测风险保障。我们将方法应用于智能代理缓存系统,构建渐进信任模型,以保证程度决定是否可自主服务缓存结果。
原文摘要 · Abstract (English)
We present a comprehensive ablation of nine finite-sample bound families for selective prediction with risk control, combining concentration inequalities (Hoeffding, Empirical Bernstein, Clopper-Pearson, Wasserstein DRO, CVaR) with multiple-testing corrections (union bound, Learn Then Test fixed-sequence) and betting-based confidence sequences (WSR). Our main theoretical contribution is Transfer-Informed Betting (TIB), which warm-starts the WSR wealth process using a source domain's risk profile, achieving tighter bounds in data-scarce settings with a formal dominance guarantee. We prove that the TIB wealth process remains a valid supermartingale under all source-target divergences, that TIB dominates standard WSR when domains match, and that no data-independent warm-start can achieve better convergence. The combination of betting-based confidence sequences, LTT monotone testing, and cross-domain transfer is, to our knowledge, a three-way novelty not present in the literature. We evaluate all nine bound families on four benchmarks-MASSIVE (n=1,102), NyayaBench (n=280), CLINC-150 (n=22.5K), and Banking77 (n=13K)-across 18 (alpha, delta) configurations. On MASSIVE at alpha=0.10, LTT eliminates the ln(K) union-bound penalty, achieving 94.0% guaranteed coverage versus 73.8% for Hoeffding-a 27% relative improvement. On NyayaBench, where the small calibration set makes Hoeffding-family bounds infeasible below alpha=0.20, Transfer-Informed Betting achieves 18.5% coverage at alpha=0.10, a 5.4x improvement over LTT + Hoeffding. We additionally compare with split-conformal prediction, showing that conformal methods produce prediction sets (avg. 1.67 classes) whereas selective prediction provides single-prediction risk guarantees. We apply these methods to agentic caching systems, formalizing a progressive trust model where the guarantee determines when cached responses can be served autonomously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。