用大模型分析5000篇论文,研究数学家如何判断证明是否具有解释力。
Using Large Language Models to Study Mathematical Practice
- 用Gemini 2.5 Pro分析arXiv上5000篇数学论文的文本
- 发现数学家频繁提及解释性,且不同领域差异明显
- 为哲学家提供非挑选的实证数据,推动实践哲学方法革新
数学实践哲学(PMP)通过实际数学工作中的证据来回答哲学问题。其中一项重要方向是研究数学解释,旨在理解数学家认为哪些证明具有解释性,以及追求解释在数学实践中扮演何种角色。为应对传统案例研究中选择性偏差和文件柜问题,一些学者近年转向语料库分析作为替代方案。本文利用Google Gemini 2.5 Pro大模型,基于5000篇arXiv论文样本,完成了大规模文本分析,获得数百个高质量标注实例。研究旨在回答:数学家在多大程度上会明确提及解释性?不同数学领域中解释性实践是否存在显著差异?哪些解释理论最符合大量非挑选的实证数据?哲学家如何进一步利用AI工具从此类大数据中获取洞见?作为首个系统使用大模型进行PMP研究的成果,本文也试图开启关于这类工具在面向实践的哲学研究中应用的讨论,并评估当前模型在此类任务中的优劣。
原文摘要 · Abstract (English)
The philosophy of mathematical practice (PMP) looks to evidence from working mathematics to help settle philosophical questions. One prominent program under the PMP banner is the study of explanation in mathematics, which aims to understand what sorts of proofs mathematicians consider explanatory and what role the pursuit of explanation plays in mathematical practice. In an effort to address worries about cherry-picked examples and file-drawer problems in PMP, a handful of authors have recently turned to corpus analysis methods as a promising alternative to small-scale case studies. This paper reports the results from such a corpus study facilitated by Google's Gemini 2.5 Pro, a model whose reasoning capabilities, advances in hallucination control and large context window allow for the accurate analysis of hundreds of pages of text per query. Based on a sample of 5000 mathematics papers from arXiv.org, the experiments yielded a dataset of hundreds of useful annotated examples. Its aim was to gain insight on questions like the following: How often do mathematicians make claims about explanation in the relevant sense? Do mathematicians' explanatory practices vary in any noticeable way by subject matter? Which philosophical theories of explanation are most consistent with a large body of non-cherry-picked examples? How might philosophers make further use of AI tools to gain insights from large datasets of this kind? As the first PMP study making extensive use of LLM methods, it also seeks to begin a conversation about these methods as research tools in practice-oriented philosophy and to evaluate the strengths and weaknesses of current models for such work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。