研究大模型时代论文语言风格是否趋同,发现中文法语仍保有母语特征。
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era

- 用半自动标注构建论文语料库,训练分类器识别作者母语痕迹。
- 跨时期分析显示母语识别率持续下降,后大模型时代出现异常趋势。
- 中文、法语表现抗同质化,日韩语则显著弱化,适合关注语言多样性的研究者。
从机器翻译到大语言模型(LLMs)的写作辅助工具演进,正在改变学术写作方式。本研究通过分析ACL Anthology论文在三个时期——神经网络前、大模型前和大模型后——的母语识别(NLI)趋势,探讨这一转变是否导致论文风格趋同。我们采用半自动化框架构建标注数据集,并微调分类器以检测作者背景的语言指纹。分析表明,母语识别性能随时间持续下降。有趣的是,大模型时代出现异常:中文和法语表现出意外的抵抗性或反常趋势,而日语和韩语则出现比预期更显著的下降。
原文摘要 · Abstract (English)
The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is homogenizing research papers by analyzing native language identification (NLI) trends in ACL Anthology papers across three eras: pre-neural network (NN), pre-LLM, and post-LLM. We construct a labeled dataset using a semi-automated framework and fine-tune a classifier to detect linguistic fingerprints of author backgrounds. Our analysis shows a consistent decline in NLI performance over time. Interestingly, the post-LLM era reveals anomalies: while Chinese and French show unexpected resistance or divergent trends, Japanese and Korean exhibit sharper-than-expected declines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。