打造高难度数学推导数据集,助力大模型逼近人类推理水平
STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework
- 通过人机协作多智能体框架生成高质量数学推导样本
- 100道最难题仅5%被顶尖模型解决,验证数据挑战性
- 适合追求数学推理能力突破的研究者使用
高质量数学数据集对提升大语言模型的推理能力至关重要。现有数据集普遍存在内容陈旧、缺乏挑战性、忽视人类思维过程、可靠性不足等问题。为此,我们提出STORM-BORN,一个源自前沿学术论文的超难数学推导数据集,包含密集的人类式近似与启发式提示。为确保质量与可靠性,我们设计了新型人机协同多智能体生成框架,融合高密度推理过滤、多智能体协作与数学专家评估。共收集2000个合成样本,并精选出100道最具挑战性问题。即使最先进的模型GPT-o1在这些题目上正确率也低于5%。基于STORM-BORN微调后,LLaMA3-8B准确率提升7.84%,Qwen2.5-7B提升9.12%。随着人工智能向数学家级推理迈进,STORM-BORN既可作为高难度基准,也可用于训练人类式推理能力。代码与数据集已公开于https://github.com/lwhere/STORM-BORN。
原文摘要 · Abstract (English)
High-quality math datasets are crucial for advancing the reasoning abilities of large language models (LLMs). However, existing datasets often suffer from three key issues: outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM generation. To address these, we introduce STORM-BORN, an ultra-challenging dataset of mathematical derivations sourced from cutting-edge academic papers, which includes dense human-like approximations and heuristic cues. To ensure the reliability and quality, we propose a novel human-in-the-loop, multi-agent data generation framework, integrating reasoning-dense filters, multi-agent collaboration, and human mathematicians' evaluations. We curated a set of 2,000 synthetic samples and deliberately selected the 100 most difficult problems. Even most advanced models like GPT-o1 solved fewer than 5% of them. Fine-tuning on STORM-BORN boosts accuracy by 7.84% (LLaMA3-8B) and 9.12% (Qwen2.5-7B). As AI approaches mathematician-level reasoning, STORM-BORN provides both a high-difficulty benchmark and a human-like reasoning training resource. Our code and dataset are publicly available at https://github.com/lwhere/STORM-BORN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。