arXiv:2602.17431cs.CLcs.AI2026-02中稿 · TMLR被引 4

为长文本生成设计细粒度不确定性量化框架,提升事实准确性。

Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study

  • 按分解、打分、聚合三阶段构建长文本不确定性评估体系
  • 发现断言-响应蕴含关系比复杂方法更有效,断言级打分优于句子级
  • 不确定性感知解码显著提升长文本事实性,适合需要高可信生成的场景

不确定性量化已成为检测大语言模型闭卷幻觉的有效手段,但现有方法多针对短文本输出,难以推广至长文本生成。本文提出一种细粒度不确定性量化的新分类体系,依据响应分解、单元级评分和响应级聚合三个阶段的设计选择进行区分。我们形式化了几类基于一致性的黑盒评分方法,对现有技术进行了泛化与扩展。同时引入FactScore-STEM-Geo,一个包含400个问题的长文本问答数据集,覆盖科学、技术、工程、数学及地理四个领域。在多个大模型和数据集上的实验表明:1)断言-响应蕴含关系在性能上始终优于或等同于更复杂的断言级评分方法;2)断言级评分整体优于句子级评分;3)不确定性感知解码能显著提升长文本生成的事实性。本框架厘清了已有方法间的关系,支持直接比较,并为组件选择提供实用指导。

原文摘要 · Abstract (English)

Uncertainty quantification has emerged as an effective approach to closed-book hallucination detection for LLMs, but existing methods are largely designed for short-form outputs and do not generalize well to long-form generation. We introduce a taxonomy for fine-grained uncertainty quantification in long-form LLM outputs that distinguishes methods by design choices at three stages: response decomposition, unit-level scoring, and response-level aggregation. We formalize several families of consistency-based black-box scorers, providing generalizations and extensions of existing methods. We also introduce FactScore-STEM-Geo, a new 400-question long-form QA dataset spanning four categories across STEM and Geography. In our experiments across multiple LLMs and datasets, we find 1) claim-response entailment consistently performs better or on par with more complex claim-level scorers, 2) claim-level scoring generally yields better results than sentence-level scoring, and 3) uncertainty-aware decoding is highly effective for improving the factuality of long-form outputs. Our framework clarifies relationships between prior methods, enables apples-to-apples comparisons, and provides practical guidance for selecting components for fine-grained UQ.

不确定性量化长文本生成事实性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。