arXiv:2607.23614cs.CRcs.AI2026-07

用验证器修复量子场论中的对偶性错误,提升大模型修复成功率。

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

  • 用符号验证器检测异常匹配、R荷一致性等四类对偶性条件。
  • 在145个错误对偶性声明上,修复成功率提升7.1至8.3个百分点。
  • 适合研究高能物理与大模型协同验证的学者使用。

我们提出DualityCert,一种用于检验四维N=1晶格规范理论中候选塞伯格对偶性声明的符号验证器。该验证器评估't Hooft反常匹配、超势R荷一致性、中心荷匹配及有界手征环代理。通过验证的声明获得一致性证书,仅表示未发现测试中的不一致,并非证明对偶性成立。我们将其作为语言模型代理的修复环境,代理接收故意破坏的对偶性声明,需持续编辑直至获得认证。在预先注册的145个错误声明基准上,验证器引导重试策略在deepseek-chat和qwen-plus上分别比单次尝试提升8.3和7.1个百分点(Holm校正p<0.002)。在相同十一轮预算下,停止优先策略组合在deepseek-chat上落后独立验证过滤重采样10.3个百分点,但在qwen-plus上领先14.7个百分点,表明两种模型下最优策略顺序反转。在qwen-plus上,类别级验证反馈带来+8.7个百分点增益,可解释义务恒等式本身贡献+6.4个百分点,而deepseek-chat未观测到类似效应。另一次预注册的MiniMax-M2.5扩展再次确认迭代收益,且独立验证过滤重采样优于策略组合。因此,最优策略因模型而异,但所有优策略均依赖同一低成本证书。验证器、基准、协议及所有尝试记录均已开源。

原文摘要 · Abstract (English)

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.

量子场论大模型修复形式验证对偶性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。