科学论文中破折号使用率在大模型时代显著上升,或反映写作方式的群体性变化。
Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era
- 分析6.9万篇medRxiv预印本,发现破折号使用率从4.23%升至11.58%
- 破折号增长呈渐进式加速,2025年已达20.3%,且经多重验证
- 破折号或为大模型影响的群体性写作痕迹,适合关注AI对学术写作影响的研究者
大型语言模型(LLMs)可能在辅助文本中留下细微的风格痕迹,其中最常被提及的是破折号(Unicode U+2014)。然而,尚无研究测量其在科学文献中的使用变化。本研究在开放科学框架(HFT8C)预注册,使用medRxiv全文XML预印本数据集,选取2020–2025年首次投稿、讨论部分可提取且不少于500字符的原始版本(N = 69,632)。主要终点为讨论部分是否存在至少一个破折号,主效应为ChatGPT前(2022年11月30日前)与后时期破折号流行率的绝对差值,采用按第一作者聚类标准误的逻辑回归模型。分析计划(六项支持分析、六项敏感性分析、两项伪造检验)在计算任何确认结果前冻结。结果显示,破折号在讨论部分的流行率从4.23%升至11.58%,绝对增加7.35个百分点(95% CI 6.94–7.77;OR 2.96,95% CI 2.77–3.17)。该增长并非突变,而是逐步加速:2023年接近4%,2024年达8.0%,2025年升至20.3%。所有可行敏感性分析均支持此效应(7.35–7.60个百分点),且两项伪造检验均未发现显著变化(预模型时代内部分组仅+0.13个百分点,95% CI -0.33 到 +0.58),且在模板段落中基本不存在。独立的与大模型相关的词汇标记及同文档章节对比也指向同一结论。破折号是群体水平的指标,非单篇论文检测大模型使用的工具,研究设计无法确立因果关系;但表明2020年代初科学写作方式发生了显著且大致同步的变化。
原文摘要 · Abstract (English)
Large language models (LLMs) can leave subtle stylistic traces in assisted text; one of the most cited is the em-dash (Unicode U+2014). Yet no one has measured whether em-dash use has changed in the scientific literature. This study, pre-registered on the Open Science Framework (HFT8C), used the full set of medRxiv full-text XML preprints from the official Text-and-Data-Mining resource. The primary cohort was first, original versions deposited 2020-2025 with an extractable Discussion section of at least 500 characters (N = 69,632). The primary endpoint was the presence of at least one em-dash in the Discussion; the principal measure was the absolute change in its prevalence between the pre-ChatGPT era (before 30 November 2022) and the post-ChatGPT era, estimated with a logistic model with standard errors clustered by first author. The analysis plan (six supporting analyses, six sensitivity analyses, two falsification tests) was frozen before any confirmatory result was computed. Em-dash prevalence in Discussion sections rose from 4.23% before ChatGPT to 11.58% afterward, an absolute increase of 7.35 percentage points (95% CI 6.94-7.77; odds ratio 2.96, 95% CI 2.77-3.17). The rise was not a sharp jump but a gradual, delayed acceleration: near 4% through 2023, 8.0% in 2024, and 20.3% in 2025. The effect survived every feasible sensitivity analysis (7.35-7.60 pp) and both falsification tests; a placebo split within the pre-LLM era showed no meaningful change (+0.13 pp, 95% CI -0.33 to +0.58), and was essentially absent in boilerplate sections. Independent LLM-associated lexical markers and within-paper section comparisons pointed the same way. The em-dash is a population-level indicator, not a per-paper detector of LLM use, and the design cannot establish causality; it shows that something in how scientific literature is written changed markedly in the early 2020s, and roughly when.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。