arXiv:2604.16217cs.CLcs.AI2026-04

用模型内部表示替代输出统计量,提升大模型问答的可靠性

Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations

论文配图:Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
图 1 · 摘自论文原文
  • 用模型各层激活信息计算非一致性得分,避免输出统计量失效
  • 跨领域场景下准确率提升显著,且在设定风险下保持可靠覆盖
  • 适合对可靠性要求高的实际部署,尤其面对分布偏移时

大语言模型在高可靠性场景中应用日益广泛,但基于输出的不确定性度量(如词元概率、熵、自一致性)在分布外或校准-部署不匹配时容易失效。置信区间预测能提供有限样本下的严格有效性,但其效果依赖于非一致性分数的质量。本文提出一种基于内部表示的置信区间预测框架,用于大模型问答任务:引入分层信息(Layer-Wise Information, LI)得分,衡量输入条件如何改变模型深度上预测熵的变化,并将其作为标准分割置信区间流程中的非一致性分数。在闭合式与开放域问答基准上,该方法在跨领域迁移下表现最优,相较于强文本级基线实现更优的有效性-效率权衡,同时在同域设置下保持与设定风险水平相当的可靠性。结果表明,当表面不确定性受分布偏移影响不稳定时,内部表示可提供更具信息量的置信评分。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in settings where reliability matters, yet output-level uncertainty signals such as token probabilities, entropy, and self-consistency can become brittle under calibration--deployment mismatch. Conformal prediction provides finite-sample validity under exchangeability, but its practical usefulness depends on the quality of the nonconformity score. We propose a conformal framework for LLM question answering that uses internal representations rather than output-facing statistics: specifically, we introduce Layer-Wise Information (LI) scores, which measure how conditioning on the input reshapes predictive entropy across model depth, and use them as nonconformity scores within a standard split conformal pipeline. Across closed-ended and open-domain QA benchmarks, with the clearest gains under cross-domain shift, our method achieves a better validity--efficiency trade-off than strong text-level baselines while maintaining competitive in-domain reliability at the same nominal risk level. These results suggest that internal representations can provide more informative conformal scores when surface-level uncertainty is unstable under distribution shift.

置信预测大模型可靠性内部表示跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。