arXiv:2602.21201cs.AIcs.CL2026-02被引 6

AI数学研究助手自主解决首届FirstProof挑战中6道题,准确率超六成。

Aletheia tackles FirstProof autonomously

  • 基于Gemini 3 Deep Think的AI自主推理求解数学问题。
  • 在10题中成功解决6题,其中第8题专家意见不统一。
  • 代码与输出公开透明,适合关注AI数学能力的研究者。

我们报告了由Gemini 3 Deep Think驱动的数学研究智能体Aletheia(Feng等,2026b)在首届FirstProof挑战中的表现。在规定时间内,Aletheia自主解决了10道题中的6道(2、5、7、8、9、10),经多数专家评估确认;仅第8题专家意见未达成一致。为确保透明,本文详述了对FirstProof的理解、实验细节及评估流程。原始提示与输出可于https://github.com/google-deepmind/superhuman/tree/main/aletheia获取。

原文摘要 · Abstract (English)

We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed timeframe of the challenge, Aletheia autonomously solved 6 problems (2, 5, 7, 8, 9, 10) out of 10 according to majority expert assessments; we note that experts were not unanimous on Problem 8 (only). For full transparency, we explain our interpretation of FirstProof and disclose details about our experiments as well as our evaluation. Raw prompts and outputs are available at https://github.com/google-deepmind/superhuman/tree/main/aletheia.

AI数学自主推理智能体数学竞赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。