arXiv:2506.15455cs.CLcs.AI2025-06ICML被引 13

构建推理能力分级评估框架,防止模型靠记忆作弊。

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

  • 基于因果三层次理论,用符号化表示生成多级推理题
  • 在4个基准上测试多个大模型,性能普遍下降
  • 适合研究模型真实推理能力的学者使用

近期大型语言模型在推理基准上报告了高准确率,但尚不清楚这些结果是源于真正推理还是训练集的统计记忆。受因果阶梯(Pearl, 2009)三层次(关联、干预、反事实)启发,本文提出RE-IMAGINE框架,用于刻画大模型推理能力的层级,并设计自动化流水线,在不同层级生成问题变体。通过在中间符号表示层面修改问题,RE-IMAGINE可生成大量仅靠记忆无法解决的新问题。该框架具有通用性,适用于数学、代码和逻辑等多个推理领域。我们在四个广泛使用的基准上对多个大模型家族进行评估,发现当模型被问及问题变体时性能普遍下降。这些结果表明模型对过往数据存在一定程度的依赖,为后续研究跨推理层级的能力打开了新路径。

原文摘要 · Abstract (English)

Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true reasoning or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (associations, interventions and counterfactuals), this paper introduces RE-IMAGINE, a framework to characterize a hierarchy of reasoning ability in LLMs, alongside an automated pipeline to generate problem variations at different levels of the hierarchy. By altering problems in an intermediate symbolic representation, RE-IMAGINE generates arbitrarily many problems that are not solvable using memorization alone. Moreover, the framework is general and can work across reasoning domains, including math, code, and logic. We demonstrate our framework on four widely-used benchmarks to evaluate several families of LLMs, and observe reductions in performance when the models are queried with problem variations. These assessments indicate a degree of reliance on statistical recall for past performance, and open the door to further research targeting skills across the reasoning hierarchy.

推理评估大模型因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。