同一问题不同问法,大模型答案常变,可靠性远低于准确率显示。
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

- 测试相同语义下不同问法对大模型输出的影响
- 23%以上问题在不同问法中出现答案翻转
- 单一正确回答不代表可靠,适合评估模型稳定性
大型语言模型在基准测试中常表现出高准确率,但其在相同问题的不同等价表述下是否稳定应用知识仍不明确。本文研究了事实问答与数学推理任务中,语义不变的改写对模型输出的影响。在四个基准和13个模型上发现,模型输出高度依赖提示词的具体措辞。尽管整体准确率变化不大,但实例级行为极不稳定:许多问题在不同问法下交替出现正确与错误答案,答案翻转率超过23%。对原题答对的问题进行分析,发现答案翻转率更高,表明单次正确不能反映可靠性。同时,多数问题至少有一个改写版本能获得正确答案,说明知识存在但检索不一致。基于此,提出简单自改写策略可部分恢复隐含知识,提升推理表现。结果表明,标准准确率可能掩盖显著不稳定性,跨等价输入的一致性评估更清晰反映大模型可靠性。
原文摘要 · Abstract (English)
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。