用心理测量学生成更难的数学题,检验大模型是否真会解题。
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
- 基于项目反应理论动态评估题目难度并生成合理变体。
- 在26个模型上使顶尖大模型得分平均下降21.73%。
- 适合评估数学推理真实能力,对抗数据泄露和模式匹配。
大型语言模型在数学推理任务中表现优异,但可能因训练数据泄露或表面模式匹配而夸大性能。为此,需要基于对抗性扰动的评测来衡量真正的数学推理能力。现有基于规则的扰动方法常生成病态问题,阻碍题目难度系统评估与基准演化。为此,我们提出RIDE,一种利用项目反应理论(IRT)严格度量题目难度并生成更难、更合理的数学问题变体的对抗性重写框架。我们使用35个大模型模拟学生,根据其答题表现构建难度排序器,作为强化学习中的奖励信号,指导问题重写模型在不同难度层级上重构题目。将RIDE应用于竞赛级数学基准,生成的扰动版本使先进大模型性能显著下降,实验显示26个模型平均得分下降21.73%,暴露了其数学推理能力的脆弱性,验证了该评测方法的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather than genuine reasoning. To this end, an adversarial perturbation-based evaluation is needed to measure true mathematical reasoning ability. Current rule-based perturbation methods often generate ill-posed questions and impede the systematic evaluation of question difficulty and the evolution of benchmarks. To bridge this gap, we propose RIDE, a novel adversarial question-rewriting framework that leverages Item Response Theory (IRT) to rigorously measure question difficulty and to generate intrinsically more challenging, well-posed variations of mathematical problems. We employ 35 LLMs to simulate students and build a difficulty ranker from their responses. This ranker provides a reward signal during reinforcement learning and guides a question-rewriting model to reformulate existing questions across difficulty levels. Applying RIDE to competition-level mathematical benchmarks yields perturbed versions that degrade advanced LLM performance, with experiments showing an average 21.73% drop across 26 models, thereby exposing limited robustness in mathematical reasoning and confirming the validity of our evaluation approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。