ML论文越写越难懂,NeurIPS该立新规了
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
- 用可量化标准评估论文可读性,发现标题缩写密度飙升
- 可读性差的论文引用少,影响传播与影响力
- 建议2027年试点缩写预算、通俗摘要等七项新规
机器学习研究呈指数级增长,但其表达规范未同步演进。本文分析1991-2025年280万篇arXiv论文、1987-2024年2.4772万篇NeurIPS论文及1990-2025年2450万篇PubMed论文,采用经典可读性评分、Hohmann写作风格套件(含煽动性语言)、缩写密度与复用率、大模型作为可读性裁判、以及OpenAlex和Semantic Scholar引文数据。结果显示:第一,NeurIPS摘要在各类可读性指标上持续恶化,Flesch阅读易度从1987年约24降至2024年13,煽动性语言在2015-2024年间上升约50%;第二,标题缩写密度从1987年每百词0.33增至2024年3.21,约89%的缩写使用次数不足十次,高于科学界平均水平十个百分点;第三,可读性更高的论文获得更多引用,表明可读性与影响力正相关,可读性差的论文可能陷入碎片化困境;第四,大模型裁判评分显示摘要可读性自1987至2022年基本稳定,2022年后出现改善迹象,此趋势与传统指标相悖,引发核心问题:目标读者是人类还是大模型?最后,NeurIPS投稿量自1987至2024年增长约50倍。若以人类读者为目标,建议2027年试点七项标准:缩写预算配官方术语表、人类可读性阈值、更严格引文标准、独立视觉元素、通俗语言摘要、预注册缩写释义表及开源审计工具。
原文摘要 · Abstract (English)
Machine learning research has grown exponentially while its communication norms have not. We argue NeurIPS should adopt explicit, measurable writing standards. We analyze 2.8 million arXiv papers (1991-2025), 24,772 NeurIPS papers (1987-2024), and 24.5 million PubMed papers (1990-2025), applying classical readability scores, the Hohmann writing style suite (including sensational language), acronym density and reuse, an LLM as judge readability protocol, and citations from OpenAlex and Semantic Scholar. Four patterns emerge. First, NeurIPS abstracts score harder to read on every classical readability metric: Flesch Reading Ease falls from about 24 in 1987 to 13 in 2024, and sensational language rises by about 50 percent in NeurIPS abstracts between 2015 and 2024. Second, acronym density in NeurIPS titles has grown from 0.33 per 100 words in 1987 to 3.21 in 2024, and about 89 percent of NeurIPS acronyms are used fewer than ten times, ten points above the science-wide baseline. Third, more readable NeurIPS papers tend to receive more citations, suggesting readability and impact are correlated and that less readable papers risk remaining fragmented. LLM as judge scores rate NeurIPS abstracts as roughly stable from 1987 to 2022, with early signs of improvement thereafter, a pattern that disagrees with every classical readability metric and raises a design question for enforcement: is the target reader a human or an LLM? Lastly, NeurIPS volume has grown roughly 50-fold between 1987 and 2024. Assuming the goal is to optimise for human readers, we propose seven standards NeurIPS could pilot at NeurIPS 2027: an acronym budget with a venue-approved term list, a human readability threshold, stricter citation standards, standalone visual elements, a plain language summary, a pre-registered acronym glossary, and open source audit tooling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。