arXiv:2605.04352cs.LGmath.GR2026-05

用代数陷阱门测试大模型的结构性数学推理能力

Probing Structural Mathematical Reasoning in Language Models with Algebraic Trapdoors

  • 构建基于SL(3,Z)子群的评测基准,具密码学式验证-求解不对称性
  • 一模型耗时152分钟,识别瓶颈并选择不回答以保持认知校准
  • 揭示模型行为四类:正确/错误承诺与正确/错误回避,超越简单评分

我们提出一个评估语言模型结构性数学推理能力的基准套件,基于SL(3, Z)中的子群构造问题,具有类似密码学的验证者-求解者不对称性。每个实例以整数矩阵列表形式给出有限生成子群,并要求计算一个算术不变量——指数、素数满射性或成员资格——该值在构造时由信息(N, K)决定为O(1)闭式表达,但求解者若无此信息,则需通过Aschbacher分类分析或在未知可判定性的SL(3, Z)中进行成员查询来推导。该基准可区分内化代数先验(如Aschbacher类、McLaughlin定理、性质(T)、同余子群性质)的模型与依赖通用计算的模型。我们在两个先进模型的五条代表性推理轨迹上报告实证结果。主要发现:在一模型的指数变体任务中,其推理耗时152分钟,明确识别出核侧成员资格问题为瓶颈,尝试构造性验证后选择不回答‘DON'T KNOW’,而非提交其计算的余核候选——展现了对开放可判定边界的认知校准,这正是本基准所设计探测的核心。我们认为该基准揭示了模型行为的四类划分(正确承诺、错误承诺、正确回避、错误回避),而传统答案键评分会混淆这些类别。

原文摘要 · Abstract (English)

We introduce a benchmark suite for evaluating structural mathematical reasoning in language models, built on subgroup-construction problems in SL(3, Z) with cryptographic-style verifier-prover asymmetry. Each instance presents a finitely generated subgroup as a list of integer matrices and asks for an arithmetic invariant -- index, surjection-at-prime, or membership -- that the construction-time information (N, K) pins down in O(1) closed form, but that the solver, lacking that information, must derive by either Aschbacher-classification analysis or by a membership query in SL(3, Z) of unknown decidability. The benchmark therefore distinguishes models with internalized algebraic priors (Aschbacher classes, McLaughlin's theorem, Property (T), the congruence subgroup property) from models that rely on general-purpose computation. We report empirical results across five representative reasoning traces from two state-of-the-art models. The headline result: on the index variant, one model spent 152 minutes of reasoning, explicitly identified the kernel-side membership question as the bottleneck, attempted constructive verification, and abstained with "DON'T KNOW" rather than commit to its computed cokernel candidate -- demonstrating calibrated meta-cognition on the open-decidability boundary that the benchmark was designed to probe. We argue that the benchmark exposes a four-way classification of model behavior (commit-correct, commit-wrong, abstain-correct, abstain-wrong) that standard answer-key scoring conflates.

数学推理代数结构模型评估认知校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。