用ChatGPT-5.2辅助证明数学猜想,验证了大模型在科研中的实用潜力
Early Evidence of Vibe-Proving with Consumer LLMs: A Case Study on Spectral Region Characterization with ChatGPT-5.2 (Thinking)
- 通过生成、评审、修复的迭代流程,用消费级大模型攻克4阶矩阵谱域猜想
- 成功给出非实谱区的充要条件及边界构造,证明结果可审计可复现
- 揭示大模型在高阶思路探索中有效,但关键纠错仍需人类专家介入
大型语言模型(LLMs)正日益成为科研协作者,但在研究级数学领域的应用证据仍有限,尤其针对个人研究者可操作的工作流。本文通过可审计的案例研究,展示了消费级订阅版LLM在“ vibe-proving”中的早期证据,解决了Ran与Teng(2024)提出的第20号猜想——关于4-循环行随机非负矩阵族的精确非实谱区域。我们分析了七条可共享的ChatGPT-5.2(Thinking)对话线程和四份版本化的证明草稿,记录了一个生成、评审、修复的迭代流程。结果显示,模型在高层次证明搜索中最为有用,而人类专家在正确性关键环节仍不可或缺。最终定理给出了必要且充分的区域条件,并提供了显式的边界达成构造。除数学成果外,本研究还从过程层面刻画了大模型辅助的实际价值边界与验证瓶颈,对评估AI辅助研究工作流及设计人机协同定理证明系统具有启示意义。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used as scientific copilots, but evidence on their role in research-level mathematics remains limited, especially for workflows accessible to individual researchers. We present early evidence for vibe-proving with a consumer subscription LLM through an auditable case study that resolves Conjecture 20 of Ran and Teng (2024) on the exact nonreal spectral region of a 4-cycle row-stochastic nonnegative matrix family. We analyze seven shareable ChatGPT-5.2 (Thinking) threads and four versioned proof drafts, documenting an iterative pipeline of generate, referee, and repair. The model is most useful for high-level proof search, while human experts remain essential for correctness-critical closure. The final theorem provides necessary and sufficient region conditions and explicit boundary attainment constructions. Beyond the mathematical result, we contribute a process-level characterization of where LLM assistance materially helps and where verification bottlenecks persist, with implications for evaluation of AI-assisted research workflows and for designing human-in-the-loop theorem proving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。