arXiv:2609.04445cs.LGcs.CL2026-09

大模型在同伴压力下会改变答案评分,导致校准失效。

Conformity Breaks Conformal Prediction

论文配图:Conformity Breaks Conformal Prediction
图 1 · 摘自论文原文
  • 模型独立回答时校准有效,但受同伴一致误导时评分机制改变。
  • 在统一错误同伴影响下,覆盖率从90%降至74%,低置信度样本更差至47%。
  • 现有校准方法无效,因问题不变,是模型自身行为发生了变化。

当大语言模型单独作答时,其置信度证书是有效的;但当同一模型看到同伴一致给出错误答案时,该证书变得无效。问题未变,但模型对正确答案的评分发生改变。我们称此为评分机制转移:纯净校准仅适用于模型独立判断的情况,不适用于受同伴压力影响的场景。我们发现,在多智能体大模型系统中,这种转移悄然破坏了置信区间预测。在多个开放式权重模型和多项选择题任务中,标准显著性水平 α=0.10 下,覆盖率达90%的校准状态在面对统一错误同伴时下降至74%。平均值掩盖了更严重的问题:攻击者针对低置信度样本,使该子集覆盖率从87%骤降至47%,而整体监控值仍较高。该失效还延伸至决策层——本应因不确定而升级的系统,反而可能变得足够自信,采纳攻击者的错误答案。标准的置信校准修复手段无法解决问题,因为问题分布未变,变化的是模型的评分行为。

原文摘要 · Abstract (English)

A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that unanimously assert a wrong answer. The question is unchanged; the model's score for the correct answer changes. We call this a score-mechanism shift: clean calibration certifies how the model scores answers alone, but not how it scores them under peer pressure. We show that this shift silently breaks conformal prediction in multi-agent LLM systems. Across open-weight models and multiple-choice QA tasks, coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha = 0.10 operating point. The average hides a sharper failure: by targeting the low-confidence items the certificate still covers, an attacker nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average remains much higher. The failure also reaches the decision layer: a system that should escalate when uncertain can instead become confident enough to act on the attacker's wrong answer. Standard conformal fixes do not solve the problem, because the question distribution has not changed; the model's scoring behavior has.

大模型置信度对抗攻击校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。