arXiv:2601.02179cs.CL2026-01ACL被引 13

首次系统研究大模型多轮对话中的置信度估计,提升可信对话能力。

Confidence Estimation for LLMs in Multi-turn Interactions

  • 提出多轮置信度评估框架,强调每轮校准与信心单调递增。
  • 新指标InfoECE显示主流方法在多轮中校准差,信心不随信息增加而上升。
  • 引入基于对数几率的探测器P(Sufficient),有效区分真实证据与闲聊内容。

尽管置信度估计是缓解大语言模型幻觉的有前景方向,但现有研究主要集中于单轮场景。多轮对话中,上下文累积、模糊性逐步消除,模型置信度动态变化仍缺乏探索。本文首次系统研究多轮交互中的置信度估计,建立基于两方面理想属性的评估框架:每轮校准性与随着信息增加信心的单调性。为支持研究,我们引入新指标,包括长度归一化的期望校准误差(InfoECE),以及用于生成受控评估数据集的“先知-猜测者”范式。实验表明,广泛使用的置信度方法在多轮对话中难以保持校准性和单调性。相比之下,我们提出的基于对数几率的探测器P(Sufficient)表现更优,能稳健追踪证据积累,区分真实信息与对话冗余。本工作为构建更可靠、可信的对话智能体提供了基础方法。

原文摘要 · Abstract (English)

While confidence estimation is a promising direction for mitigating hallucinations in Large Language Models (LLMs), current research overwhelmingly focuses on single-turn settings. The dynamics of model confidence in multi-turn conversations, where context accumulates and ambiguity is progressively resolved, remain largely unexplored. This work presents the first systematic study of confidence estimation in multi-turn interactions, establishing a formal evaluation framework grounded in two key desiderata: per-turn calibration and monotonicity of confidence as more information becomes available. To facilitate this, we introduce novel metrics, including a length-normalized Expected Calibration Error (InfoECE), and a new "Hinter-Guesser" paradigm for generating controlled evaluation datasets. Our experiments reveal that widely-used confidence techniques struggle with calibration and monotonicity in multi-turn dialogues. In contrast, a novel logit-based probe we introduce, P(Sufficient), proves comparatively more effective, robustly tracking evidence accumulation and distinguishing it from conversational filler. Our work provides a foundational methodology for developing more reliable and trustworthy conversational agents.

置信度估计大模型多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。