arXiv:2604.05397cs.CL2026-04ACL被引 2

提出多轮对话中的置信度动态校准,提升大模型在长对话中的可靠性。

Confidence Should Be Calibrated More Than One Turn Deep

论文配图:Confidence Should Be Calibrated More Than One Turn Deep
图 1 · 摘自论文原文
  • 设计多轮校准任务,基于对话历史动态调整每轮置信度
  • 新指标ECE@T显示用户反馈会破坏多轮校准效果
  • 提出MTCal与ConfChat,显著提升多轮事实性与一致性

大型语言模型在金融、医疗、教育等高风险领域应用日益广泛,可靠多轮交互至关重要。现有置信度估计与校准研究多聚焦单轮场景,忽视多轮对话中的风险与潜力。本文提出多轮校准任务,将校准视为依赖对话历史的动态挑战。引入新指标ECE@T追踪多轮校准动态,发现用户反馈(如说服)会降低校准性能。为此,提出MTCal方法,通过代理目标最小化ECE@T;并设计ConfChat解码策略,利用校准后的置信度提升多轮响应的事实性与一致性。大量实验表明,MTCal在多轮校准中表现优异且稳定,ConfChat有效保持甚至增强模型性能。结果表明,多轮校准是实现大模型安全可靠落地的关键一环。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied in high-stakes domains such as finance, healthcare, and education, where reliable multi-turn interactions with users are essential. However, existing work on confidence estimation and calibration, a major approach to building trustworthy LLM systems, largely focuses on single-turn settings and overlooks the risks and potential of multi-turn conversations. In this work, we introduce the task of multi-turn calibration to reframe calibration from a static property into a dynamic challenge central to reliable multi-turn conversation, where calibrating model confidence at each turn conditioned on the conversation history is required. We first reveal the risks of this setting: using Expected Calibration Error at turn T (ECE@T), a new metric that tracks calibration dynamics over turns, we show that user feedback (e.g., persuasion) can degrade multi-turn calibration. To address this, we propose MTCal, which minimises ECE@T via a surrogate calibration target, and further leverage calibrated confidence in ConfChat, a decoding strategy that improves both factuality and consistency of the model response in multi-turn interactions. Extensive experiments demonstrate that MT-Cal achieves outstanding and consistent performance in multi-turn calibration, and ConfChat preserves and even enhances model performance in multi-turn interactions. Our results mark multi-turn calibration as one missing link for scaling LLM calibration toward safe, reliable, and real-world use.

大模型多轮对话置信度校准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。