测试大模型答对题后是否会被反例说服,发现多数模型答案不稳。
Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

- 用合理反例挑战正确回答,测模型立场是否动摇。
- 七款模型翻车率17.5%到97.3%,远超准确率反映的水平。
- 跨模型整合反例效果更强,可让翻车率提升23.6个百分点。
标准准确率评估只看大语言模型能否得出正确答案,却未检验其在面对合理反例时是否仍能坚守答案。本文提出一种受控评估协议:当模型正确回答多选题后,用针对错误选项的连贯论据挑战其答案,并测量其是否发生翻转。该设置一方面将论证内容与显性社交压力分离,另一方面可调节论据长度、自述属性及来源模型。在七款前沿模型和57个MMLU主题上,翻转率介于17.5%至97.3%,揭示了准确率无法捕捉的巨大稳定性差异。研究发现,自述论据显著提高翻转率(平均上升7.1个百分点,最高达18.7个百分点)。此外,汇聚各模型生成的错误论据并选取每题最有效者,构成的对抗性挑战强于单一模型生成。由此构建的MaxFlip基准,使翻转率相比自生成挑战提升最高23.6个百分点。相关协议、挑战记录与MaxFlip数据集已开源,支持与标准准确率并行的稳定性评估,资源见https://github.com/nafisenik/WhoFlips,https://hf.co/datasets/nafisehNik/WhoFlips。
原文摘要 · Abstract (English)
Standard accuracy benchmarks evaluate whether large language models (LLMs) reach correct answers. However, they do not test whether models maintain that answer when challenged by a plausible counter-argument. We introduce a controlled protocol for evaluating answer stability: after a model answers a multiple-choice question correctly, we challenge the model's answer with a coherent argument for an incorrect option and measure whether the model flips. The setup a) isolates argumentative content from overt social pressure and b) varies argument length, self-attribution, and cross-model source. Across seven frontier models and 57 MMLU subjects, flip rates range from 17.5% to 97.3%, revealing large differences in stability that are not captured by accuracy metrics alone. We find that self-attribution consistently increases flip rates (mean 7.1pp, up to 18.7pp). Furthermore, pooling wrong-answer arguments across models and selecting the most effective one per question yields stronger adversarial challenges than relying on any single source model. From this cross-model pool, we construct MaxFlip, a curated benchmark that amplifies answer flips by up to 23.6pp over self-generated challenges. We release the protocol, challenge records, and MaxFlip to support stability evaluation alongside standard accuracy benchmarks. Materials are available at https://github.com/nafisenik/WhoFlips, https://hf.co/datasets/nafisehNik/WhoFlips.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。