LLM生成形式化证明时,轻微语义改写会导致输出大幅波动。
Evaluating Autoformalization Robustness via Semantically Similar Paraphrasing
- 用语义相似的改写句测试LLM形式化能力,评估其鲁棒性
- 在MiniF2F和ProofNet数据集上,改写后准确率下降超30%
- 提示工程师需警惕自然语言微调对形式化结果的影响
大型语言模型(LLMs)在自动形式化任务中表现强劲,但仍可能生成缺乏依据且难以验证的形式化内容。近期文本到SQL研究发现,即使自然语言输入语义高度一致,LLMs对改写仍敏感。本文将此现象引入形式化领域,通过生成语义相似的改写自然语言句,评估两个现代LLM在MiniF2F和Lean 4版ProofNet基准上的形式化输出鲁棒性,测量语义有效性和编译有效性。结果表明,自然语言的细微变化会显著影响模型输出,性能在不同改写句间存在明显波动。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently emerged as powerful tools for autoformalization. Despite their impressive performance, these models can still struggle to produce grounded and verifiable formalizations. Recent work in text-to-SQL, has revealed that LLMs can be sensitive to paraphrased natural language (NL) inputs, even when high degrees of semantic fidelity are preserved. In this paper, we investigate this claim in the autoformalization domain. Specifically, we evaluate the robustness of LLMs generating formal proofs with semantically similar paraphrased NL statements by measuring semantic and compilation validity. Using the formal benchmarks MiniF2F and Lean 4 version of ProofNet, and two modern LLMs, we generate paraphrased natural language statements and cross-evaluate these statements across both models. The results of this paper reveal performance variability across paraphrased inputs, demonstrating that minor shifts in NL statements can significantly impact model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。