arXiv:2603.03330cs.CLcs.AI2026-03被引 3

测试大模型在质疑下的稳定性,发现部分模型会无理由改答案

Certainty robustness: Evaluating LLM stability under self-challenging prompts

  • 设计双轮问答框架,模拟用户质疑场景
  • 200道题测试中,30%模型正确答案被错误修改
  • 强调模型自信度与回答一致性,适合评估真实对话可靠性

大语言模型常以高置信度输出答案,但缺乏对自身确定性的显式推理机制。现有基准多关注单轮准确率、真实性或置信度校准,未涵盖交互场景下模型面对质疑时的表现。我们提出Certainty Robustness Benchmark,一个两轮评估框架,通过不确定性提问(如“你确定吗?”)和直接反驳(如“你错了!”)等自挑战提示,结合数值置信度获取,评估模型在交互中的稳定性和适应性。基于LiveBench的200道推理与数学题,我们评估了四款前沿LLM,区分了合理自修正与不合理答案变更。结果揭示:模型间交互可靠性差异显著,仅靠基线准确率无法解释;部分模型在对话压力下放弃正确答案,而另一些则表现出强抗质疑能力,且置信度与正确性更一致。这表明,确定性鲁棒性是衡量模型对齐性、可信度及实际部署价值的关键新维度。

原文摘要 · Abstract (English)

Large language models (LLMs) often present answers with high apparent confidence despite lacking an explicit mechanism for reasoning about certainty or truth. While existing benchmarks primarily evaluate single-turn accuracy, truthfulness or confidence calibration, they do not capture how models behave when their responses are challenged in interactive settings. We introduce the Certainty Robustness Benchmark, a two-turn evaluation framework that measures how LLMs balance stability and adaptability under self-challenging prompts such as uncertainty ("Are you sure?") and explicit contradiction ("You are wrong!"), alongside numeric confidence elicitation. Using 200 reasoning and mathematics questions from LiveBench, we evaluate four state-of-the-art LLMs and distinguish between justified self-corrections and unjustified answer changes. Our results reveal substantial differences in interactive reliability that are not explained by baseline accuracy alone: some models abandon correct answers under conversational pressure, while others demonstrate strong resistance to challenge and better alignment between confidence and correctness. These findings identify certainty robustness as a distinct and critical dimension of LLM evaluation, with important implications for alignment, trustworthiness and real-world deployment.

模型评估可靠性交互鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。