用AI找出大模型数学短板,自动生成难题测试其真实能力
Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis
- 基于AI假设分析大模型错题,定位具体薄弱知识点
- 生成难题使Llama-3.3-70B准确率降至45%,远低于原基准77%
- 可扩展至其他领域,适合评估模型泛化能力
现有数学评测基准多依赖人工构建,难以规模化且易导致模型过拟合。尽管已有自动化生成方法,但多数未聚焦大模型的具体错误原因,且仅限特定类别。为此,我们提出一种新管道:利用AI生成假设,识别大模型在数学概念与技能上的薄弱点,并针对性生成新题目。实验表明,假设准确度越高,生成题目越难——基于最准确假设生成的问题使Llama-3.3-70B-Instruct准确率降至45%,相较原MATH基准的77%显著下降。该方法高度灵活,可拓展至非数学领域,为评估大模型跨域表现提供有力工具。
原文摘要 · Abstract (English)
Numerous math benchmarks exist to evaluate LLMs' mathematical capabilities. However, most involve extensive manual effort and are difficult to scale. Consequently, they cannot keep pace with LLM development or easily provide new instances to mitigate overfitting. Some researchers have proposed automatic benchmark generation methods, but few focus on identifying the specific math concepts and skills on which LLMs are error-prone, and most can only generate category-specific benchmarks. To address these limitations, we propose a new math benchmark generation pipeline that uses AI-generated hypotheses to identify the specific math concepts and skills that LLMs struggle with, and then generates new benchmark problems targeting these weaknesses. Experiments show that hypothesis accuracy positively correlates with the difficulty of the generated problems: problems generated from the most accurate hypotheses reduce Llama-3.3-70B-Instruct's accuracy to as low as 45%, compared to 77% on the original MATH benchmark. Furthermore, our pipeline is highly adaptable and can be applied beyond math to explore a wide range of LLM capabilities, making it a valuable tool for investigating how LLMs perform across different domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。