arXiv:2512.13764cs.AIcs.LG2025-12

数学与编程是评估AI能力的通用基准,能推动AI自我进化。

Mathematics and Coding are Universal AI Benchmarks

  • 用形式化证明系统构建数学纤维,实现稳定自提升
  • 数学与编程任务组合可覆盖全部评估空间,编码更通用
  • 适合研究通用智能评估与自迭代AI系统的学者

我们研究数学与编程在人工智能代理心理测评体系模空间中的特殊作用。基于先前工作的AAI框架与GVU动力学,定义了数学纤维,并发现当结合形式化证明内核(如Lean、Coq)时,该纤维上的GVU流具有谱稳定的自提升特性,源于类似预言机的验证机制。主要技术成果为一个密度定理:在代理输出均匀紧致且AAI泛函满足Lipschitz条件下,由数学定理证明与编程任务生成的测评子空间,在评估度量下稠密于整个测评模空间。仅编程本身具备普遍性,纯数学则不具备;其优势在于谱特性而非表达能力。这一结果表明,数学与编程构成评估的“通用坐标”,而形式数学是高级AI实现递归自提升的自然起点。

原文摘要 · Abstract (English)

We study the special role of mathematics and coding inside the moduli space of psychometric batteries for AI agents. Building on the AAI framework and GVU dynamics from previous works, we define the Mathematics Fiber and show that, when paired with formal proof kernels (e.g. Lean, Coq), GVU flows on this fiber admit spectrally stable self-improvement regimes due to oracle-like verification. Our main technical result is a density theorem: under uniform tightness of agent outputs and a Lipschitz AAI functional, the subspace of batteries generated by mathematical theorem-proving and coding tasks is dense in the moduli space of batteries with respect to the evaluation metric. Coding alone is universal in this sense, while pure mathematics is not; its privilege is spectral rather than expressive. We interpret this as evidence that mathematics and coding provide ``universal coordinates'' for evaluation, and that formal mathematics is a natural ignition domain for recursive self-improvement in advanced AI agents.

AI评估自提升形式证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。