arXiv:2606.09449cs.CL2026-06

无需标准答案,用分轴检查实现自动形式化系统的自我修正。

Reasoning without Gold Standards: A Proxy-Judge Theory of Autoformalization

论文配图:Reasoning without Gold Standards: A Proxy-Judge Theory of Autoformalization
图 1 · 摘自论文原文
  • 用多维度结构化代理判断替代标准答案匹配
  • 在多个数据集上使通过率超越单次提示基线
  • 适合无可靠参考答案的复杂推理任务研究者

复杂推理任务常需评估模型输出的正确性,但无法依赖单一参考答案进行精确匹配。自动形式化(AF)即为代表:将非正式数学或逻辑推理转化为可验证的形式对象,然而专家标注的形式化难以扩展至真实案例,且同一非正式论证可能对应多种有效形式表达。因此,进展依赖于能否用部分、结构化的代理判断替代精确参考。本文提出一种无参考的代理法官框架,以多轴属性检查取代黄金标准匹配,涵盖全局结构、模块内部及跨域对齐三个层次,并聚合为判断向量。该向量驱动反思式精炼循环,违反项引导控制器修复对应目标,每次仅调整错误部分。在有界判断噪声下,期望内在差距几何收敛至与噪声相关的平台值。在 miniF2F、ProofNet、e-SNLI 与 ProntoQA 上,七种形式化骨干模型均显示精炼显著提升通过率,且分轴代理优于标量代理,尤其在基线仍有改进空间时。结构化代理判断既提供实用精炼信号,也为无标准参考场景下的收敛性提供了理论保障。

原文摘要 · Abstract (English)

Complex reasoning tasks increasingly require systems to produce outputs whose correctness cannot be judged by exact match against a single reference. Autoformalization (AF) is a representative example; it asks a model to translate informal mathematical or logical reasoning into a formally checkable object, yet expert-validated formalizations do not scale beyond toy cases and a single informal argument can admit many valid formal renderings. Progress therefore depends on whether partial, structured proxies can substitute for exact references. We introduce a reference-free proxy-judge framework for AF that replaces gold-standard matching with a vector of per-axis property checks. The framework organizes the proxy along three structural scopes that cover global properties of the elicited object, per-module properties internal to its sub-components, and cross-domain properties that re-align it to the informal source, and aggregates each axis into a verdict vector. The vector drives a reflective refinement loop in which a violated coordinate routes the controller to a matching repair target, so each iteration changes only what is judged wrong. Under bounded judge noise, the expected intrinsic gap contracts geometrically to a noise-dependent plateau. Across seven formalization backbones on miniF2F, ProofNet, e-SNLI, and ProntoQA, refinement consistently lifts Pass Rate over the single-shot ICL baseline, and the per-axis proxy outperforms a matched scalar proxy on benchmarks where the baseline has room to improve. Structured proxy judgments therefore provide both a practical refinement signal and a theoretical handle on convergence when exact references are unavailable.

自动形式化自我修正代理判断推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。