通过追溯训练数据,揭示大模型自信表达的真正来源。
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
- 基于信息检索与影响估计,追踪模型自信表达的训练数据源头。
- 发现OLMo2-13B常受无关词汇影响,依赖表面表达而非内容真实依据。
- 适合关注大模型可信度、可解释性与训练机制的研究者阅读。
大语言模型(LLMs)通过输出自信表述可提升用户信任感,但现有研究表明,其自信常被高估,与事实准确性不一致。为理解这种自信表达的来源,我们提出TracVC(Tracing Verbalized Confidence),结合信息检索与影响估计,将模型生成的自信表达回溯至训练数据。在OLMo和Llama模型的问答任务中评估,引入新指标“内容根基性”(content groundness),衡量模型是否基于与问题/答案相关的内容训练例建立自信,而非仅依赖通用自信表达样本。分析显示,OLMo2-13B频繁受与查询语义无关的自信相关数据影响,可能仅模仿表层语言形式表达确定性,而非真实内容支撑。这暴露了当前训练范式的核心缺陷:模型学会如何表现自信,却未掌握何时该自信。该研究为提升大模型自信表达的可靠性提供了基础。
原文摘要 · Abstract (English)
Large language models (LLMs) can increase users' perceived trust by verbalizing confidence in their outputs. However, prior work has shown that LLMs are often overconfident, making their stated confidence unreliable since it does not consistently align with factual accuracy. To better understand the sources of this verbalized confidence, we introduce TracVC (\textbf{Trac}ing \textbf{V}erbalized \textbf{C}onfidence), a method that builds on information retrieval and influence estimation to trace generated confidence expressions back to the training data. We evaluate TracVC on OLMo and Llama models in a question answering setting, proposing a new metric, content groundness, which measures the extent to which an LLM grounds its confidence in content-related training examples (relevant to the question and answer) versus in generic examples of confidence verbalization. Our analysis reveals that OLMo2-13B is frequently influenced by confidence-related data that is lexically unrelated to the query, suggesting that it may mimic superficial linguistic expressions of certainty rather than rely on genuine content grounding. These findings point to a fundamental limitation in current training regimes: LLMs may learn how to sound confident without learning when confidence is justified. Our analysis provides a foundation for improving LLMs' trustworthiness in expressing more reliable confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。