arXiv:2502.15795cs.AIcs.CL2025-02被引 5

用高质量合成数据提升数学证明自动形式化,少样本更有效。

Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization

  • 通过精心设计的回译提示生成高精度数学证明数据
  • 仅用1/150令牌量超越多语言微调模型在ProofNet上的表现
  • 适合资源有限但追求高精度数学推理的AI研究者

自动形式化——将非正式数学语言转化为形式化规范与证明——仍是当前大语言模型面临的难题。现有研究对性能差距存在不同解释。为此,我们提出一种新方法:利用人工优化提示进行回译,以增强语言模型的数学能力,尤其解决标注数据稀缺问题。具体评估三种策略:(1) 实时(在线)回译,(2) 离线提炼回译结合少样本增强,(3) 结合证明状态信息的逐行分析。各方案均强调数据质量而非数量,聚焦生成证明的高保真度。结果表明,采用所提方法生成的合成数据能显著提升LLM在ProofNet等标准基准上的自动形式化表现。关键发现:仅用极少量令牌(仅为MMA多语言数据集的1/150),即超越预训练模型微调效果。本方法为降低数学证明形式化资源消耗提供了新路径,推动数学AI发展。

原文摘要 · Abstract (English)

Autoformalization, the process of transforming informal mathematical language into formal specifications and proofs remains a difficult task for state-of-the-art (large) language models. Existing works point to competing explanations for the performance gap. To this end, we introduce a novel methodology that leverages back-translation with hand-curated prompts to enhance the mathematical capabilities of language models, particularly addressing the challenge posed by the scarcity of labeled data. Specifically, we evaluate three primary variations of this strategy: (1) on-the-fly (online) backtranslation, (2) distilled (offline) backtranslation with few-shot amplification, and (3) line-by-line proof analysis integrated with proof state information. Each variant is designed to optimize data quality over quantity, focusing on the high fidelity of generated proofs rather than sheer data scale. Our findings provide evidence that employing our proposed approaches to generate synthetic data, which prioritizes quality over volume, improves the Autoformalization performance of LLMs as measured by standard benchmarks such as ProofNet. Crucially, our approach outperforms pretrained models using a minimal number of tokens. We also show, through strategic prompting and backtranslation, that our approaches surpass the performance of fine-tuning with extensive multilingual datasets such as MMA on ProofNet with only 1/150th of the tokens. Taken together, our methods show a promising new approach to significantly reduce the resources required to formalize proofs, thereby accelerating AI for math.

自动形式化数学AI数据质量小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。