arXiv:2601.23166cs.CL2026-01被引 1

无需参考答案,自动优化定理形式化质量

Monotonic Reference-Free Refinement for Autoformalization

  • 推理时迭代优化,不依赖真值或人工干预
  • 多维度联合提升,达100%形式有效性
  • 适合自动化证明与数学语言转换场景

尽管命题形式化已取得快速进展,但完整定理的形式化仍鲜有探索。现有迭代优化方法通常仅改进语法正确性等单一维度,难以同时优化多个质量指标,而这对完整定理形式化至关重要。本文提出一种推理时的无参考、单调迭代过程,利用定理证明器和基于大模型的评判者提供互补反馈,无需真实标注或已有形式化结果,也无需人工介入。该方法在形式有效性、逻辑保真性、数学一致性与形式质量的掩码复合目标上进行优化,由响应图指导不同角色大模型优先改善特定维度。我们还设计了一种保证认证单调提升的接受策略,并给出收敛与终止条件。实验表明,该方法可同时提升多个维度,在miniF2F上实现100.00%形式有效性与90.27%综合得分,在ProofNet上达到77.96%形式有效性与52.45%综合得分。

原文摘要 · Abstract (English)

While statement autoformalization has advanced rapidly, full-theorem autoformalization remains largely unexplored. Existing iterative refinement methods in statement autoformalization typically improve isolated aspects of formalization, such as syntactic correctness, but struggle to jointly optimize multiple quality dimensions, which is critical for full-theorem autoformalization. We introduce a reference-free iterative monotonic process at inference time for full-theorem autoformalization that leverages complementary feedback from theorem provers and LLM-based judges, without access to ground-truth or existing formalizations and without human intervention. Our approach optimizes a masked composite objective over Formal Validity, Logical Preservation, Mathematical Consistency, and Formal Quality, guided by a responsiveness map that indicates how different LLMs acting as different roles preferentially improve each dimension. We further propose an acceptance policy that guarantees certified monotonic improvement, and provide conditions ensuring convergence and termination. Empirical experiments demonstrate the proposed process enables simultaneous improvement across multiple dimensions, achieving 100.00% formal validity and a 90.27% overall score on miniF2F, and 77.96% formal validity and a 52.45% overall score on ProofNet.

形式化大模型自动证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。