arXiv:2603.01239cs.CLcs.AI2026-03被引 1

多轮对话中大模型自信会自发漂移,影响判断可靠性。

Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence

  • 通过自锚定机制分析模型在多轮对话中的自信变化趋势。
  • 不同模型表现各异:部分模型自信下降,部分则上升或抑制改善。
  • 揭示了模型自我迭代时校准能力退化的潜在风险,适合可信AI研究者关注。

我们提出自锚定校准漂移(SACD),即大语言模型在多轮对话中基于自身历史输出逐步构建时,表达的自信出现系统性变化的假设。通过对比三种前沿模型(Claude Sonnet 4.6、Gemini 3.1 Pro、GPT-5.2)在150个涵盖事实、技术与开放问题领域的测试中,在单轮基线(A)、多轮自锚定(B)和独立重复控制(C)三种条件下进行实证研究。结果显示复杂且模型异质:Claude Sonnet 4.6在自锚定下自信显著降低(均值CDS = -0.032,t(14) = -2.43,p = .029,d = -0.627),同时校准误差漂移显著(F(4,56) = 22.77,p < .001,eta² = .791)。GPT-5.2在开放领域呈现相反趋势(均值CDS = +0.026),且在第5轮时ECE显著上升。Gemini 3.1 Pro未显示显著CDS(t(14) = 0.38,p = .710),但其条件C数据表明:无自锚定时,其校准误差从0.327降至接近零;而自锚定使其保持在约0.333——说明SACD可能表现为对自然校准提升的抑制而非主动恶化。

原文摘要 · Abstract (English)

We introduce Self-Anchoring Calibration Drift (SACD), a hypothesized tendency for large language models (LLMs) to show systematic changes in expressed confidence when building iteratively on their own prior outputs across multi-turn conversations. We report an empirical study comparing three frontier models -- Claude Sonnet 4.6, Gemini 3.1 Pro, and GPT-5.2 -- across 150 questions spanning factual, technical, and open-ended domains, using three conditions: single-turn baseline (A), multi-turn self-anchoring (B), and independent repetition control (C). Results reveal a complex, model-heterogeneous pattern that partially diverges from pre-registered hypotheses. Claude Sonnet 4.6 exhibited significant decreasing confidence under self-anchoring (mean CDS = -0.032, t(14) = -2.43, p = .029, d = -0.627), while also showing significant calibration error drift (F(4,56) = 22.77, p < .001, eta^2 = .791). GPT-5.2 showed the opposite pattern in open-ended domains (mean CDS = +0.026) with significant ECE escalation by Turn 5. Gemini 3.1 Pro showed no significant CDS (t(14) = 0.38, p = .710), but its Condition C data reveals a striking ECE pattern: without self-anchoring, Gemini's calibration error drops from .327 to near zero across repetitions, whereas self-anchoring holds ECE flat at approximately .333 -- indicating that SACD can manifest as suppression of natural calibration improvement rather than ac

大模型校准自信漂移多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。