通过多维度响应分析,更精准识别大模型回答的不确定性。
Uncertainty Quantification of Large Language Models through Multi-Dimensional Responses
- 生成多个响应,结合语义与知识维度相似性分析
- 在医疗等高风险场景中,不确定响应识别准确率显著提升
- 适合关注模型可靠性与可信度的研究者与应用开发者
大型语言模型(LLMs)因大规模训练数据和强大的Transformer架构,在各类任务中表现出色。然而,其输出结果的可靠性仍存疑。在医疗、金融及决策等领域,对LLM进行不确定性量化(UQ)至关重要。现有方法主要关注语义相似性,忽略了响应中隐含的知识维度。本文提出一种多维UQ框架,融合语义与知识感知的相似性分析。通过生成多个响应,并利用辅助LLM提取隐含知识,构建独立的相似性矩阵,再通过张量分解获得综合不确定性表征。该方法分离了语义与知识维度的重叠信息,同时捕捉语义变化与事实一致性,实现更精确的不确定性评估。实验表明,该方法在识别不确定响应方面优于现有技术,为高风险应用中的模型可靠性提供了更稳健的保障。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks due to large training datasets and powerful transformer architecture. However, the reliability of responses from LLMs remains a question. Uncertainty quantification (UQ) of LLMs is crucial for ensuring their reliability, especially in areas such as healthcare, finance, and decision-making. Existing UQ methods primarily focus on semantic similarity, overlooking the deeper knowledge dimensions embedded in responses. We introduce a multi-dimensional UQ framework that integrates semantic and knowledge-aware similarity analysis. By generating multiple responses and leveraging auxiliary LLMs to extract implicit knowledge, we construct separate similarity matrices and apply tensor decomposition to derive a comprehensive uncertainty representation. This approach disentangles overlapping information from both semantic and knowledge dimensions, capturing both semantic variations and factual consistency, leading to more accurate UQ. Our empirical evaluations demonstrate that our method outperforms existing techniques in identifying uncertain responses, offering a more robust framework for enhancing LLM reliability in high-stakes applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。