arXiv:2607.16868cs.AI2026-07

用逻辑图建模答案间蕴含与矛盾,更准判断大模型输出可信度。

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

  • 构建逻辑图捕捉答案间的蕴含与不相容关系。
  • 在逻辑结构化问题上,不确定性评估准确率提升7.1%(AUROC)。
  • 适合需要精准识别幻觉的高风险场景应用。

大语言模型常给出看似自信却不可靠的输出,给安全敏感场景部署带来挑战。现有不确定性度量如语义熵仅关注语义等价性,忽略不同答案间的逻辑关系,导致过度估计不确定性,并在形式多样但逻辑兼容的回答中误判为幻觉(如粒度或具体程度差异)。本文提出逻辑图不确定性(LGU)框架,显式建模答案间的蕴含与不相容关系。LGU将概率质量沿蕴含链聚集到最具体的假设上,计算该分布的熵,并惩罚假设间的相互不相容性。在多个问答基准和模型族上,LGU在现有度量中平均排名第一,尤其在逻辑结构化问题上,相较于语义熵,最高提升7.1% AUROC和3.5% AUARC。

原文摘要 · Abstract (English)

Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains onto the most specific hypotheses the answers support, measures the entropy of the resulting distribution, and penalizes mutual incompatibility among those hypotheses. Across multiple question-answering benchmarks and model families, LGU ranks first on average among existing uncertainty measures, with its largest gains---up to +7.1\% AUROC and +3.5\% AUARC over semantic entropy---on questions whose sampled answers are logically structured.

不确定性逻辑推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。