arXiv:2506.05038cs.CL2025-06中稿 · Findings of ACL202…被引 4

自动测试大模型数学推理的脆弱性,发现隐藏缺陷并提升抗干扰能力。

Toward Automated Robustness Evaluation of Mathematical Reasoning

  • 通过多轮重写验证生成对抗性题目,保持语义一致同时诱导模型出错。
  • 在GSM8K和MATH-500上成功暴露模型漏洞,失败率提升显著。
  • 可扩展至非数学任务,且生成数据能有效提升模型鲁棒性。

大型语言模型在各类推理任务中表现出色,但对相同任务的简单变体常表现出意外脆弱性。现有鲁棒性评估主要依赖人工设计模板或有限扰动规则,缺乏适应性,易受数据污染影响。为此,我们提出数学压力测试框架MaSTer,受软件压力测试启发,通过多轮重写-验证循环生成对抗性题目,在保证语义一致性的同时有效引发模型失败。该框架为每种LLM动态生成基准变体,显著降低数据污染风险。在GSM8K和MATH-500上的实验表明,MaSTer能有效检测数学推理中的脆弱性。此外,我们验证了其在非数学任务上的可扩展性,并证明由MaSTer生成的变体可用于微调,显著提升模型鲁棒性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpected brittleness, often failing on simple variations of the same underlying task. Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules. Consequently, such approaches lack the adaptability to probe latent vulnerabilities unique to specific models and remain susceptible to data contamination. To address this, we propose the Math Stress Tester (MaSTer), an automated framework inspired by software stress testing. MaSTer generates adversarial variants via a multi-round rewrite-verify loop, ensuring semantic consistency while successfully inducing model failure. Our framework generates benchmark variants dynamically for each LLM, thus minimizing the risk of data contamination. Experiments on GSM8K and MATH-500 demonstrate the effectiveness of MaSTer on mathematical tasks. Additionally, we validate the framework's extensibility to non-mathematical tasks, highlighting its broad applicability. Furthermore, we demonstrate that the synthesized variants generated by MaSTer can be utilized as a fine-tuning dataset to significantly enhance the model's robustness.

大模型评测数学推理鲁棒性测试对抗生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。