arXiv:2506.23864cs.CL2025-06Conference of the …被引 5

三大推理基准存在严重设计缺陷,模型高分未必真会推理。

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It

  • 用大模型检测基准题目的结构、语义和表述问题
  • 清洗后测试发现分数提升源于表面变化而非真实推理能力
  • 适合关注模型评估可信度的研究者和开发者

我们系统性审计了三个广泛使用的推理基准:SocialIQa、FauxPas-EAI 和 ToMi,发现其题目设计和评估方法存在普遍性缺陷。利用五种大语言模型(GPT-{3, 3.5, 4, o1} 和 LLaMA 3.1)作为诊断工具,我们识别出题目中存在重复项、歧义表述和不合逻辑答案等问题,以及评分机制过度关注输出形式而非推理过程。通过人工标注与清理后的子集重新评估,发现模型得分提升主要来自表面语句微调,而非真正推理能力的增强。进一步分析表明,模型表现对输入细节(如上下文是否提供、措辞方式)极为敏感,高分可能反映对格式线索的迎合,而非基于信息的一致推断。该研究质疑当前基于基准的推理能力声称的有效性,呼吁采用更注重推理过程的评估协议。我们已公开经审计的数据与评估工具,以支持更具可解释性的模型推理诊断。

原文摘要 · Abstract (English)

We conduct a systematic audit of three widely used reasoning benchmarks, SocialIQa, FauxPas-EAI, and ToMi, and uncover pervasive flaws in both benchmark items and evaluation methodology. Using five LLMs (GPT-{3, 3.5, 4, o1}, and LLaMA 3.1) as diagnostic tools, we identify structural, semantic, and pragmatic issues in benchmark design (e.g., duplicated items, ambiguous wording, and implausible answers), as well as scoring procedures that prioritize output form over reasoning process. Through systematic human annotation and re-evaluation on cleaned benchmark subsets, we find that model scores often improve not due to due to erratic surface wording variations and not to improved reasoning. Infact, further analyses show that model performance is highly sensitive to minor input variations such as context availability and phrasing, revealing that high scores may reflect alignment with format-specific cues rather than consistent inference based on the input. These findings challenge the validity of current benchmark-based claims about reasoning in LLMs, and highlight the need for evaluation protocols that assess reasoning as a process of drawing inference from available information, rather than as static output selection. We release audited data and evaluation tools to support more interpretable and diagnostic assessments of model reasoning.

模型评估推理能力基准审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。