arXiv:2509.12527cs.LGstat.ML2025-09被引 6

用信息提升统计量为大模型输出提供可靠置信度,实现高精度错误拦截。

Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design

  • 通过对比模型概率与骨架基线,利用子伽马 PAC-Bayes 边界积累证据。
  • 在8个数据集上实现2%风险下77%覆盖率,优于基线平均10个百分点。
  • 适合对安全性要求高的场景,如医疗、金融等高风险应用。

大语言模型常产生自信但错误的输出,亟需具备正式弃权保证的可靠不确定性量化方法。本文提出信息提升证书,通过比较模型概率与骨架基线,利用子伽马 PAC-Bayes 边界累积证据,在重尾分布下仍保持有效性,克服了标准浓度不等式的局限。在八个不同数据集上,该方法在2%风险下达到77.0%覆盖率,较近期基线平均提升10.0个百分点。在高风险场景中,相较熵基方法(18-31%拦截率),可阻断96%的关键错误。尽管频率基认证无法保证严重性加权安全且依赖骨架质量,其性能在分布偏移下仍能平滑退化,适用于真实部署。

原文摘要 · Abstract (English)

Large language models often produce confident but incorrect outputs, creating a critical need for reliable uncertainty quantification with formal abstention guarantees. We introduce information-lift certificates that compare model probabilities to a skeleton baseline, accumulating evidence through sub-gamma PAC-Bayes bounds that remain valid under heavy-tailed distributions where standard concentration inequalities fail. On eight diverse datasets, our method achieves 77.0\% coverage at 2\% risk, outperforming recent baselines by 10.0 percentage points on average. In high-stakes scenarios, we block 96\% of critical errors compared to 18-31\% for entropy-based methods. While our frequency-based certification does not guarantee severity-weighted safety and depends on skeleton quality, performance degrades gracefully under distributional shifts, making the approach practical for real-world deployment.

不确定性量化大模型安全置信度评估PAC-Bayes

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。