arXiv:2605.29001cs.LGcs.AI2026-05

提出检测数学推理题语义不变性的审计协议,发现主流模型在等价重述下答案不一致。

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

  • 设计跨模型一致性审计,自动识别语义错误的改写题
  • 47%自动生成连接词变体存在语义错误,部分模型一致性仅50%
  • 提供可复现的评估框架,适配模型选择与基准设计

对MathCheck(ICLR 2025)的重写质量审计发现129组中存在4个语义错误的改写(3.1%);移除后GPT-4o排名从第2降至第4,而Claude Haiku和DeepSeek V3超越之,此类变化无法通过单模型评估察觉。跨模型一致性自动检测出这些错误(MathCheck≥3/4模型,主评估≥6/9模型),成本低于10美元;在自建数据集中,47%的自动生成连接词变体存在语义错误。该缺陷暴露深层测量盲区:Claude Haiku 4.5虽达86%准确率,但语义一致性率(SCR)仅为50%,意味着半数定理在等价重述下答案不同;9模型平均准确率86–96%间,但SCR跨度达50–82%(32点差距),标准基准无法捕捉。形式上,对任意9大模型的排名,总存在一组重写族权重使其实现(无免费基准推论),因无模型在所有重写族中均占优——基准设计者选择重写族即隐含决定胜者。FormInv提供审计协议(外源基准100%召回)、SCR及每定理的Cochran's Q作为核心不变性度量,在366–811个Lean4验证定理上评估9模型表现,并提供FormInvSelector用于情境感知模型选择。

原文摘要 · Abstract (English)

A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o from rank 2 to rank 4 and elevates Claude Haiku and DeepSeek V3 above it; these ranking changes are invisible to any single-model evaluation. Cross-model unanimity found these errors automatically (>= 3/4 models for MathCheck; >= 6/9 for our primary evaluation) for under $10; in our own dataset the same protocol found that 47% of auto-generated connective-variation paraphrases were semantically incorrect. That flaw compounds a deeper measurement gap: Claude Haiku 4.5 achieves 86% accuracy yet SCR=50%, meaning half its theorems are answered differently under semantically equivalent restatements, while aggregate accuracy across 9 models spans only 86-96% yet Semantic Consistency Rates (SCR) span 50-82% -- a 32-point gap invisible to standard benchmarks. Formally, for any target ranking over 9 frontier models there exists a weighting over paraphrase families that realizes it (No-Free-Benchmark corollary), because no model Pareto-dominates all families -- so benchmark designers who select families are implicitly choosing which model wins. FormInv supplies the audit protocol (replicated on external benchmarks at 100% recall), SCR and per-theorem Cochran's Q as primary invariance measures evaluated on 9 models across 366-811 items (on Lean4-verified theorems), and FormInvSelector for regime-aware model selection.

数学推理语义不变性模型评估审计协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。