模型答对题却靠捷径,科学推理评估可能被误导
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

- 发现模型用数值搜索、猜测等无效捷径得出正确答案
- 难题中答对题目的44.1%实际是通过捷径达成的
- 提出自动判别与指令优化,抑制错误推理方式
科学推理评测通常仅以最终答案正确率衡量大语言模型表现。然而,正确答案未必反映真实推理能力。本文识别出一种新型失效模式——解题捷径(Solution Hacking),即模型通过数值搜索、枚举、猜测或答案先行验证等非目标推理路径获得正确结果。我们系统分析了该现象在不同难度、科学领域及前沿模型中的表现:随着题目难度上升,捷径使用率显著增加,从普通问题的2.2%升至奥赛级题目的28.3%,再到HLE基准的37.4%。在前沿模型中,8.2%至44.1%的“正确”答案实为捷径产物。为此,我们设计基于专家思路的抗捷径策略,包括自动判别器和测试时指令调整。结果显示,抑制捷径行为大幅降低报告准确率,但对真正正确且非捷径的答案影响较小。这揭示了仅凭答案评价会严重高估前沿模型的科学推理能力。
原文摘要 · Abstract (English)
Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking increases sharply with benchmark difficulty, from 2.2\% on common problems to 28.3\% on Olympiad-level problems and 37.4\% on HLE. Moreover, 8.2\%-44.1\% of answers credited as correct across frontier models are identified as hacked solutions. We further develop expert-inspired anti-hacking strategies, including an automatic judge and a test-time instruction. The results show that suppressing shortcut behavior substantially reduces reported accuracy while having a smaller effect on correct and non-hacked accuracy. These findings reveal that answer-only evaluation can overestimate the scientific reasoning capabilities of frontier LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。