让AI写数学公式时能自我检查并修正错误,提升准确性。
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
- 引入自我反思机制,逐步优化形式化表达
- 在4个基准上比最强基线提升22.6个百分点
- 适合需要高精度数学形式化的研究者使用
自动形式化将自然语言数学问题转化为机器可验证的形式化陈述,对利用形式化推理解决自然语言数学问题至关重要。尽管大语言模型能生成语法正确的形式化陈述,但常无法保持原问题的语义意图。这一局限源于现有方法将自动形式化视为简单翻译任务,缺乏人类专家常用的自我反思与迭代修正机制。为此,我们提出ReForm,一种将语义一致性评估紧密集成到自动形式化过程中的反思式方法。该方法使模型能够迭代生成形式化陈述,评估其语义保真度,并通过渐进式修正识别出的错误。为有效训练此反思模型,我们引入前瞻性有界序列优化(PBSO),在不同序列位置采用不同奖励,确保模型同时具备准确的形式化能力和正确的语义验证能力,防止表面化评价破坏反思目的。在四个自动形式化基准上的广泛实验表明,ReForm平均比最强基线提升22.6个百分点。为进一步确保评估可靠性,我们引入ConsistencyCheck,一个包含859个专家标注项的基准,不仅验证了大语言模型作为评判者的可行性,还揭示了自动形式化的内在难度:即使人类专家也在高达38.5%的情况下产生语义错误。
原文摘要 · Abstract (English)
Autoformalization, which translates natural language mathematics into machine-verifiable formal statements, is critical for using formal mathematical reasoning to solve math problems stated in natural language. While Large Language Models can generate syntactically correct formal statements, they often fail to preserve the original problem's semantic intent. This limitation arises from the LLM approaches' treating autoformalization as a simplistic translation task which lacks mechanisms for self-reflection and iterative refinement that human experts naturally employ. To address these issues, we propose ReForm, a Reflective Autoformalization method that tightly integrates semantic consistency evaluation into the autoformalization process. This enables the model to iteratively generate formal statements, assess its semantic fidelity, and self-correct identified errors through progressive refinement. To effectively train this reflective model, we introduce Prospective Bounded Sequence Optimization (PBSO), which employs different rewards at different sequence positions to ensure that the model develops both accurate autoformalization and correct semantic validations, preventing superficial critiques that would undermine the purpose of reflection. Extensive experiments across four autoformalization benchmarks demonstrate that ReForm achieves an average improvement of 22.6 percentage points over the strongest baselines. To further ensure evaluation reliability, we introduce ConsistencyCheck, a benchmark of 859 expert-annotated items that not only validates LLMs as judges but also reveals that autoformalization is inherently difficult: even human experts produce semantic errors in up to 38.5% of cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。