AI助研系统Albilich集成计算机代数,可自主解决数学难题并生成可复现的证明。
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration
- 通过整合CAS、文献检索与持久化上下文,实现长程数学推理的可控协调。
- 在RealMath上10题全解,对群论开放问题给出反例和强化证明。
- 支持人类干预,显著降低推理消耗,适合数学研究者与AI协作探索。
大型语言模型可在数学研究中提供有益思路,但长周期证明任务仍面临协调难、评估难、复现难的问题。我们提出Albilich,一个开源的数学自主研究代理框架,融合长程推理、计算机代数系统(CAS)、文献检索与基于SQLite的持久化上下文管理。在RealMath基准(Zhang et al. 2025)和Kourovka笔记中的群论开放问题(Khukhro and Mazurov 2026)上进行评估:使用CAS时在RealMath上解决10/10题,不使用时解决9/10题;在Kourovka问题中,为第21.142题构造反例,为第20.2题提供强化证明。对问题17.91的消融实验显示,启用CAS后令牌消耗降低32.0%。对问题21.142的分析表明,缺少顾问代理时验证拒绝率升高且无法合成证明路径。结果表明,Albilich是一个可由人类引导、依赖CAS增强的可扩展数学研究辅助环境。
原文摘要 · Abstract (English)
Large language models can contribute useful ideas to mathematical research, yet long-horizon proof attempts remain difficult to coordinate, evaluate, and reproduce. We present Albilich, an open-source agentic harness for autoresearch in mathematics that combines long-horizon reasoning, computer algebra systems (CAS), literature retrieval, and persistent SQLite-based context management. We evaluate Albilich on the RealMath benchmark (Zhang et al. 2025) and on open problems in group theory from the Kourovka Notebook (Khukhro and Mazurov 2026). It solved 10/10 problems on RealMath with CAS and 9/10 with no CAS. On the Kourovka problems, Albilich produced a counterexample to Problem 21.142 and a proof of a strengthening of Problem20.2. Anablation on Problem 17.91 demonstrates 32.0% token reduction when CAS is enabled. An ablation on Problem 21.142 demonstrates higher verifier-rejection rate and failure to synthesize proof routes in the absence of the advisor agent. These results support Albilich as a human-steerable, CAS-boosted environment for scalable AI-assisted mathematical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。