arXiv:2510.08615cs.CL2025-10

用大模型迭代生成数学题干扰项,保持答案不变且无需重写解法。

Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems

  • 设计迭代提示框架,从多角度生成干扰条件。
  • 确保干扰项不改变原题答案,省去重写解法的繁琐工作。
  • 适合用于构建高质量数学推理评测数据集。

数学推理是检验大语言模型智能的重要基准,数学应用题(MWPs)是常见题型。现有数据集多仅包含必要信息,干扰项问题常被忽略。先前研究发现,引入干扰项后模型性能显著下降,但现有含干扰项的数据集数量有限,且难度偏低、表达不合语境,导致干扰项易被识别排除,削弱了评测可信度。此外,添加干扰项可能改变题目推理路径和答案,需大量人工校验与重写解法。为此,本文设计一种基于LLM的迭代生成与优化框架,通过多角度提示引导生成干扰项,并建议后续修改方向。关键优势在于:新旧题目共享同一解法,通过显式指令确保干扰项不改变原始答案,避免重复生成解题过程。该方法高效易部署,有效降低构造高质干扰题的成本,同时保障数据质量。

原文摘要 · Abstract (English)

Mathematical reasoning serves as a crucial testbed for the intelligence of large language models (LLMs), and math word problems (MWPs) are a popular type of math problems. Most MWP datasets consist of problems containing only the necessary information, while problems with distracting and excessive conditions are often overlooked. Prior works have tested popular LLMs and found a dramatic performance drop in the presence of distracting conditions. However, datasets of MWPs with distracting conditions are limited, and most suffer from lower levels of difficulty and out-of-context expressions. This makes distracting conditions easy to identify and exclude, thus reducing the credibility of benchmarking on them. Moreover, when adding distracting conditions, the reasoning and answers may also change, requiring intensive labor to check and write the solutions. To address these issues, we design an iterative framework to generate distracting conditions using LLMs. We develop a set of prompts to revise MWPs from different perspectives and cognitive levels, encouraging the generation of distracting conditions as well as suggestions for further revision. Another advantage is the shared solutions between original and revised problems: we explicitly guide the LLMs to generate distracting conditions that do not alter the original solutions, thus avoiding the need to generate new solutions. This framework is efficient and easy to deploy, reducing the overhead of generating MWPs with distracting conditions while maintaining data quality.

数学推理干扰项生成LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。