美国国会新闻稿中破折号使用率在2025年翻倍,或反映大模型写作的普及。
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
- 分析480个国会办公室14.6万份新闻稿,检测无空格破折号密度变化。
- 2025年无空格破折号密度达0.217/千字符,是2021-2024年均值的两倍以上。
- 结果提示大模型辅助写作可能已广泛渗透,但不构成因果结论。
大型语言模型(LLMs)可能在文本中留下细微风格痕迹,其中最常讨论的是破折号(U+2014),特别是无空格形式的 word---word。这种用法在排版英语中常见,但在美国新闻写作中罕见,因美联社(AP)风格要求使用空格。本研究通过预注册设计(OSF: 10.17605/OSF.IO/U5NEY),分析了2021-2025年间来自480个众议院与参议院办公室的146,239份抓取自国会新闻稿的数据(开放国会新闻数据集)。以每千字符清洁文本中无空格破折号的密度为指标,采用泊松/负二项分布模型并加入长度偏移量,按办公室聚类。2021-2024年密度稳定在0.10–0.12之间,2025年跃升至0.217,超过此前四年的基线两倍;含此类破折号的稿件比例从约13%增至24.8%。主要频率比(2023-2025比2021-2022)为1.55(95% CI 1.28–1.93;预设阈值1.528),略高于1.5倍标准。该上升为新增现象(连字符密度稳定),在75.6%的262个持续办公单位中出现(p ~ 1e-16),且在224个封闭面板中持续存在。虚假检验显示:三个假阳性截断点均无效,2024/2025边界无突变,持续办公单位仍呈现上升趋势。分段回归未发现ChatGPT发布时点的突变,但发现后期加速;2025年上升在两党及两院间对称。由于预注册验证门限被正式突破,完整预注册决策规则未满足,因此解释(大模型辅助写作随模型成熟而扩散)为探索性。破折号仅作为群体层面标志,不能用于单篇文档作者识别,研究亦不支持因果推断。
原文摘要 · Abstract (English)
Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ratio (2023-2025 vs 2021-2022) was 1.55 (95% CI 1.28-1.93; exact registered cut-off: 1.528), just above the prespecified 1.5x threshold. The rise was net-new (hyphen density stable), held within authors (75.6% of 262 continuous offices increased; p ~ 1e-16) and in a closed panel of 224 offices, and survived falsification tests: three placebo cut-offs were null, the pipeline showed no step at the 2024/2025 boundary, and continuing offices carried the rise. A segmented regression finds no step at the ChatGPT cut-off but a clear post-period acceleration; the 2025 rise is symmetric across parties and chambers. Because the registered validation gate was formally breached, the full preregistered decision rule was not met; the interpretation (broad diffusion of LLM-assisted writing as the models matured) is offered as exploratory. The em-dash remains a population-level marker, not a per-release authorship detector, and the design supports no causal claim.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。