用词元熵构建可信生成的不确定性量化方法
TECP: Token-Entropy Conformal Prediction for LLMs
- 以词元熵为无参考的不确定性度量,融入分治合符预测框架
- 在六个大模型上实现可靠覆盖且预测集更紧凑
- 适合黑箱场景下需要可信赖生成的开发者使用
开放域语言生成的不确定性量化仍是关键但研究不足的挑战,尤其在无法访问内部模型信号的黑箱环境下。本文提出词元熵合符预测(TECP)框架,利用词元级别熵作为无需逻辑输出、无需参考的不确定性度量,并将其集成到分治合符预测(CP)流程中,构建具有严格覆盖率保证的预测集。与依赖语义一致性启发式或白盒特征的方法不同,TECP直接从采样生成的词元熵结构中估计认知不确定性,并通过CP分位数校准不确定性阈值,实现可证明的误差控制。在六个大型语言模型和两个基准(CoQA与TriviaQA)上的实证评估表明,TECP始终具备可靠的覆盖率和紧凑的预测集,优于先前基于自洽性的不确定性量化方法。该方法为黑箱大模型环境下的可信生成提供了原理严谨且高效的解决方案。
原文摘要 · Abstract (English)
Uncertainty quantification (UQ) for open-ended language generation remains a critical yet underexplored challenge, especially under black-box constraints where internal model signals are inaccessible. In this paper, we introduce Token-Entropy Conformal Prediction (TECP), a novel framework that leverages token-level entropy as a logit-free, reference-free uncertainty measure and integrates it into a split conformal prediction (CP) pipeline to construct prediction sets with formal coverage guarantees. Unlike existing approaches that rely on semantic consistency heuristics or white-box features, TECP directly estimates epistemic uncertainty from the token entropy structure of sampled generations and calibrates uncertainty thresholds via CP quantiles to ensure provable error control. Empirical evaluations across six large language models and two benchmarks (CoQA and TriviaQA) demonstrate that TECP consistently achieves reliable coverage and compact prediction sets, outperforming prior self-consistency-based UQ methods. Our method provides a principled and efficient solution for trustworthy generation in black-box LLM settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。