人机协作发现组合设计新定理,给出拉丁方不平衡性的紧下界。
Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design
- 大模型+符号计算+人类引导的神经符号协同推理
- 首次获得n≡1(mod 3)时拉丁方不平衡性下界4n(n−1)/9
- 适合对数学发现、人机协作研究感兴趣的学者
我们从神经符号推理的角度研究数学发现,一个由大型语言模型(LLM)驱动的AI代理,结合符号计算工具与人类战略指导,共同在组合设计理论中取得新成果。主要成果是针对难解情形n ≡ 1 (mod 3)的拉丁方不平衡性给出了紧下界。通过多日多次会话的交互日志重建发现过程,识别出各组件的独特贡献:AI有效揭示隐藏结构并生成假设;符号组件(计算机代数、约束求解器、模拟退火)实现严格验证与穷举枚举;人类引导实现了关键研究转向,将死胡同变为有效探索。分析表明,前沿大模型在批评与错误检测上可靠,但在建设性主张上不可靠。最终成果——下界4n(n−1)/9——通过一类新型近完美置换实现,并在Lean 4中完成形式化验证。实验表明,神经符号系统确实能在纯数学中产生真正发现。
原文摘要 · Abstract (English)
We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result in combinatorial design theory. The main result of this human-AI collaboration is a tight lower bound on the imbalance of Latin squares for the notoriously difficult case $n \equiv 1 \pmod{3}$. We reconstruct the discovery process from detailed interaction logs spanning multiple sessions over several days and identify the distinct cognitive contributions of each component. The AI agent proved effective at uncovering hidden structure and generating hypotheses. The symbolic component consists of computer algebra, constraint solvers, and simulated annealing, which provides rigorous verification and exhaustive enumeration. Human steering supplied the critical research pivot that transformed a dead end into a productive inquiry. Our analysis reveals that multi-model deliberation among frontier LLMs proved reliable for criticism and error detection but unreliable for constructive claims. The resulting human-AI mathematical contribution, a tight lower bound of $4n(n{-}1)/9$, is achieved via a novel class of near-perfect permutations. The bound was formally verified in Lean 4. Our experiments show that neurosymbolic systems can indeed produce genuine discoveries in pure mathematics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。