提出自适应粒度与语义聚类方法,提升长文本生成的可信度评估效率。
AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text Generation

- 用NLI中性概率区分无关与不确定内容,减少无效计算
- 基于GMM软聚类识别主题并加权聚合,提升可靠性评估精度
- 在BIO和LongFact数据集上速度提升60%,相关性达当前最佳
大语言模型在长文本生成中表现卓越,但幻觉问题严重影响其应用。不确定性量化(UQ)对评估可靠性至关重要,然而复杂结构导致跨异质主题的可靠聚合困难,且现有方法常忽略中性信息,细粒度分解又带来高昂计算成本。为此,我们提出专为长文本生成设计的UQ框架AGSC(自适应粒度与基于高斯混合模型的语义聚类)。AGSC首先利用自然语言推理(NLI)中性概率作为触发器,区分无关内容与不确定性,降低冗余计算;随后采用高斯混合模型(GMM)进行软聚类,建模潜在语义主题,并为下游聚合分配主题感知权重。在BIO和LongFact数据集上的实验表明,AGSC在事实性相关性上达到当前最优水平,同时相比全原子分解方法将推理时间减少了约60%。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in long-form generation, yet their application is hindered by the hallucination problem. While Uncertainty Quantification (UQ) is essential for assessing reliability, the complex structure makes reliable aggregation across heterogeneous themes difficult, in addition, existing methods often overlook the nuance of neutral information and suffer from the high computational cost of fine-grained decomposition. To address these challenges, we propose AGSC (Adaptive Granularity and GMM-based Semantic Clustering), a UQ framework tailored for long-form generation. AGSC first uses NLI neutral probabilities as triggers to distinguish irrelevance from uncertainty, reducing unnecessary computation. It then applies Gaussian Mixture Model (GMM) soft clustering to model latent semantic themes and assign topic-aware weights for downstream aggregation. Experiments on BIO and LongFact show that AGSC achieves state-of-the-art correlation with factuality while reducing inference time by about 60% compared to full atomic decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。