梳理大模型不确定性量化方法,揭示幻觉成因与检测路径
A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions
- 构建分类体系统一不同不确定性评估方法
- 揭示大模型幻觉与其响应置信度的关联性
- 适合关注AI可信性与安全性的研究者参考
大型语言模型在内容生成、编程和常识推理方面表现出色,已广泛应用于社会多个领域。然而,其容易产生看似合理却事实错误的幻觉内容,且表达高度自信,引发对可靠性和可信度的担忧。研究表明,通过分析模型对特定提示的不确定性,可有效检测幻觉等非事实性输出,推动了大模型不确定性量化研究的发展。本文系统综述现有不确定性量化方法,梳理其核心特征、优缺点,并构建相关分类体系,统一看似分散的方法以促进理解。同时探讨该技术在聊天机器人、文本应用及机器人具身智能中的实际应用。最后指出当前开放挑战,旨在激发未来研究。
原文摘要 · Abstract (English)
The remarkable performance of large language models (LLMs) in content generation, coding, and common-sense reasoning has spurred widespread integration into many facets of society. However, integration of LLMs raises valid questions on their reliability and trustworthiness, given their propensity to generate hallucinations: plausible, factually-incorrect responses, which are expressed with striking confidence. Previous work has shown that hallucinations and other non-factual responses generated by LLMs can be detected by examining the uncertainty of the LLM in its response to the pertinent prompt, driving significant research efforts devoted to quantifying the uncertainty of LLMs. This survey seeks to provide an extensive review of existing uncertainty quantification methods for LLMs, identifying their salient features, along with their strengths and weaknesses. We present existing methods within a relevant taxonomy, unifying ostensibly disparate methods to aid understanding of the state of the art. Furthermore, we highlight applications of uncertainty quantification methods for LLMs, spanning chatbot and textual applications to embodied artificial intelligence applications in robotics. We conclude with open research challenges in uncertainty quantification of LLMs, seeking to motivate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。