arXiv:2510.05962cs.AIcs.CL2025-10被引 1

用动态变换构造防记忆数学测评,揭示模型真实推理能力。

MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization

  • 测试时动态生成题目,数字和符号含义随机改变
  • 模型推理能力在诱导任务中表现差,但回归标准数学
  • 适合评估模型是否真会推理,而非记忆答案

数学能力评估常受污染困扰:模型可能在公开测试集后记忆答案,且现有基准题型符号与规则多样性不足,封闭答案易导致过拟合。本文提出利用这些缺陷构建动态反事实基准——MatheMagic,通过随机改变数字与运算符的语义生成新题,答案仍可自动验证。题目在测试时随机生成,用于评估模型的归纳与演绎能力,具备稳定性、可扩展性、可比性及抗过拟合优势。实验发现模型更擅长演绎而非归纳,但最终仍回归标准数学。进一步分析表明,数学适应模型缺乏通用推理能力,且在归纳任务上微调效果差。

原文摘要 · Abstract (English)

Conducting contamination-free evaluation of mathematical capabilities can be difficult for two reasons: models may memorize a test set once it is made public, and current mathematical benchmarks are prone to overfitting due to having limited diversity of symbols and rules, coupled with closed-ended answers. This paper proposes a method to leverage these shortcomings as useful features to a construct dynamic, counterfactual benchmark, which can be used to both reveal overfitting and measure true reasoning. We demonstrate this via MatheMagic, which generates math test instances with the interpretations of numbers and operators altered, yet has automatically verifiable answers. Test instances are randomly seeded and constructed at test time to evaluate a model's induction or deduction capability, offering stability, extensibility, comparability, and robustness to overfitting. Our experiments find that models solve deduction more easily than induction, but they revert to standard math. Further analysis reveals that math-adapted models fail to exhibit a general "skill" of reasoning, and fine-tuning on induction tasks generalizes poorly.

数学推理动态评测抗过拟合归纳推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。