arXiv:2608.07540cs.AI2026-08中稿 · 28th International…

测试大模型能否识别数学定理的等价变形,发现其知识很脆弱。

TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

论文配图:TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
图 1 · 摘自论文原文
  • 通过变换定理条件的数学形式,检验模型对等价表达的识别能力。
  • 最佳模型仅在60.73%的情况下正确还原定理身份,多数出错。
  • 适合研究形式化知识鲁棒性或数学AI可解释性的研究人员。

AI系统在灵活输入与正式数学对象间运行时,面临一个关键挑战:如何识别陌生表述是否对应已知的正式对象。本文以定理识别为切入点,研究在保持等价的前提下,对定理条件进行公式级变换后,模型能否恢复其对应的标准化定理身份。为此提出TREAT基准,不依赖文本改写,而是改变定理条件的数学形式,如残差方程、见证陈述、优化恒等式、集合关系、算子形式及证明中间特征。基于爬取的定理页面,筛选出可用数学表达式,提取标准定理条件,并生成带假设记录和逆映射的变体。最终数据集包含737个定理身份和29,480个转换样本。在测试集上,最佳模型仅在60.73%的案例中成功检索到正确定理身份,其他系统则表现出拒绝回答、误检或输出格式错误等不同失败模式。这表明定理知识在等价表示变化下极为脆弱。TREAT为此类形式化知识的鲁棒性评估提供受控测试平台,对需稳定目标对象、显式等价关系、验证流程与可审计评分的领域具有广泛意义。

原文摘要 · Abstract (English)

AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an equivalence-preserving transformation of a theorem condition, a model must recover the theorem identity associated with the standard statement. We introduce TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving formula-level transformations. Rather than paraphrasing theorem text, TREAT changes the mathematical form of theorem conditions themselves, expressing known results through residual equations, witness statements, optimization identities, set relations, operator forms, and proof-intermediate characterizations. Starting from scraped theorem pages, we filter for entries with usable mathematical expression forms, extract canonical theorem conditions, and generate transformed variants with recorded assumptions and inverse mappings. The final corpus contains 737 theorem identities and 29,480 transformed rows. On a test panel, the best model retrieves the correct theorem identity in only 60.73% of cases. Other systems reveal different failure modes, including abstention, wrong detection, and malformed outputs. These suggest that theorem knowledge can be fragile under equivalent changes in representation. TREAT therefore provides a controlled testbed for evaluating representation-robust access to formal knowledge, with broader relevance to domains that require stable target objects, explicit equivalence relations, validation procedures, and auditable scoring.

定理识别形式知识数学AI鲁棒性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。