arXiv:2412.17758cs.CLcs.AI2024-12被引 2

评测方式误导了对大模型能力的判断,公平评估能显著降低性能差距。

In Case You Missed It: ARC 'Challenge' Is Not That Challenging

  • 用更合理的评测方法可避免因选择项无法直接比较导致的误判。
  • 在SIQA上性能差距大幅缩小,OpenBookQA甚至出现超人表现。
  • 适合关注评测规范与模型真实能力评估的研究者参考。

ARC Challenge 对现代大模型而言看似比 ARC Easy 更难,实则源于评测设置限制了答案选项的直接比较,而非题目本身复杂度。尽管部分研究者一年来已悄然采用更合理的评估方案,但其影响尚未被广泛认知。本文揭示这一被忽视的转变,说明类似评测方式会错误暗示其他基准存在推理缺陷,并证明公平方法能显著缩小性能差距(如在SIQA上),甚至在OpenBookQA上实现超人表现。由此揭示评测设计如何影响对模型难度的感知,提出确保多选题评估真实反映模型能力的指导原则。

原文摘要 · Abstract (English)

ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. Although some researchers have quietly shifted to a more appropriate scheme over the last year, the implications of this change have yet to be widely acknowledged. We highlight this overlooked shift, show how similar evaluation practices falsely imply reasoning deficits in other benchmarks, and demonstrate that fairer methods dramatically reduce performance gaps (e.g. on SIQA) and even yield superhuman results (OpenBookQA). In doing so, we reveal how evaluation shapes perceived difficulty and offer guidelines to ensure that multiple-choice evaluations accurately reflect actual model capabilities.

大模型评测评估偏差多选题能力误判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。