arXiv:2511.10543math.HOcs.AI2025-11

用自动化系统分析3.7万篇论文,发现数学错误无处不在,连大师作品也难幸免。

From Euler to Today: Universal Mathematical Fallibility A Large-Scale Computational Analysis of Errors in ArXiv Papers

  • 构建可自动检测数学错误并生成评审报告的系统
  • 数值分析领域错误率达9.6%,几何拓扑为6.5%
  • 能评估论文适合发表的期刊等级,适用于多学科

我们对arXiv库中超过3.7万篇数学论文进行了大规模计算分析,构建了一个不仅能检测数学错误,还能生成完整评审报告并推荐期刊级别的自动化系统。该系统覆盖多个数学领域,揭示了显著的错误率和质量分布。令人惊讶的是,系统在跨越三个世纪的论文中均发现了错误,包括欧拉(1707–1783)和狄利克雷(1805–1859)的作品,以及当代菲尔兹奖得主的论文。在数值分析(math.NA)领域,错误率为9.6%(2,271例错误,共23,761篇),几何拓扑(math.GT)为6.5%(862例错误,共13,209篇)。而范畴论(math.CT)在93篇分析样本中未发现错误,暗示其更易被自动化分析。除错误检测外,系统还评估论文的期刊适配性:0.4%适合顶级综合期刊,15.5%适合顶级领域期刊,其余归类至专业刊物。这些结果表明数学错误具有普遍性,且大规模自动化同行评审在技术上可行。该方法虽聚焦数学,但具有跨学科适用性,可拓展至物理、计算机科学等arXiv涵盖领域。

原文摘要 · Abstract (English)

We present the results of a large-scale computational analysis of mathematical papers from the ArXiv repository, demonstrating a comprehensive system that not only detects mathematical errors but provides complete referee reports with journal tier recommendations. Our automated analysis system processed over 37,000 papers across multiple mathematical categories, revealing significant error rates and quality distributions. Remarkably, the system identified errors in papers spanning three centuries of mathematics, including works by Leonhard Euler (1707-1783) and Peter Gustav Lejeune Dirichlet (1805-1859), as well as contemporary Fields medalists. In Numerical Analysis (math.NA), we observed an error rate of 9.6\% (2,271 errors in 23,761 papers), while Geometric Topology (math.GT) showed 6.5\% (862 errors in 13,209 papers). Strikingly, Category Theory (math.CT) showed 0\% errors in 93 papers analyzed, with evidence suggesting these results are ``easier'' for automated analysis. Beyond error detection, the system evaluated papers for journal suitability, recommending 0.4\% for top generalist journals, 15.5\% for top field-specific journals, and categorizing the remainder across specialist venues. These findings demonstrate both the universality of mathematical error across all eras and the feasibility of automated comprehensive mathematical peer review at scale. This work demonstrates that the methodology, while applied here to mathematics, is discipline-agnostic and could be readily extended to physics, computer science, and other fields represented in the ArXiv repository.

数学错误自动化评审arXiv分析学术质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。