arXiv:2412.15184cs.LG2024-12被引 11

现有数学推理数据集过于简单且只关注结果,无法训练真正懂推理的AI助手。

Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning

  • 用更贴近真实科研过程的'有动机的证明'构建数据集
  • 当前基准测试已因过度优化而失去评估价值
  • 适合开发能理解证明思路的AI数学助手的研究者

当前用于训练和评估基于AI的数学助手机器学习模型的数据集与评测标准存在诸多缺陷,涵盖数学复杂性范围狭窄、无法捕捉证明背后的动机或思维过程等。这些问题在类似古德哈特定律的动态下加剧:当模型优化目标变为提升基准分数时,基准本身反而弱化了对真实数学能力的衡量。本文系统分析这些局限,并主张必须重构数学数据集设计与模型评估标准。应从仅提供定理到证明的直接映射,转向支持证明过程与发现过程监督的新型数据集。特别提出借鉴1949年波利亚提出的'有动机的证明'概念,作为构建更优学习信号的数据集蓝图,以缓解现有问题。

原文摘要 · Abstract (English)

The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several shortcomings and misdirections. These range from a restricted scope of mathematical complexity to limited fidelity in capturing aspects beyond the final, written proof (e.g. motivating the proof, or representing the thought processes leading to a proof). These issues are compounded by a dynamic reminiscent of Goodhart's law: as benchmark performance becomes the primary target for model development, the benchmarks themselves become less reliable indicators of genuine mathematical capability. We systematically explore these limitations and contend that enhancing the capabilities of large language models, or any forthcoming advancements in AI-based mathematical assistants (copilots or ``thought partners''), necessitates a course correction both in the design of mathematical datasets and the evaluation criteria of the models' mathematical ability. In particular, it is necessary for benchmarks to move beyond the existing result-based datasets that map theorem statements directly to proofs, and instead focus on datasets that translate the richer facets of mathematical research practice into data that LLMs can learn from. This includes benchmarks that supervise the proving process and the proof discovery process itself, and we advocate for mathematical dataset developers to consider the concept of "motivated proof", introduced by G. Pólya in 1949, which can serve as a blueprint for datasets that offer a better proof learning signal, alleviating some of the mentioned limitations.

数学推理数据集设计提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。