分析论文中LLM使用痕迹,发现用词模式正在悄然变化。
Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers
- 通过线性模型分析学术论文用词变化,揭示LLM影响
- 标题中'beyond'和'via'出现频率上升,'the'减少
- 适合关注AI对学术写作影响的研究者阅读
通过对arXiv论文的分析,我们报告了若干可能由大语言模型(LLMs)引发但此前未受充分关注的词汇使用变化,例如标题中' beyond'和' via'的频率增加,以及摘要中' the'和' of'的频率下降。由于不同LLM之间的相似性,实验表明当前分类器在多类分类任务中难以准确判断文本由哪个具体模型生成。同时,模型间差异也导致学术论文中的词汇使用模式持续演变。通过采用直接且高度可解释的线性方法,并考虑模型与提示间的差异,我们定量评估了这些影响,结果表明真实世界中LLM的使用具有异质性和动态性。
原文摘要 · Abstract (English)
Through an analysis of arXiv papers, we report several shifts in word usage that are likely driven by large language models (LLMs) but have not previously received sufficient attention, such as the increased frequency of "beyond" and "via" in titles and the decreased frequency of "the" and "of" in abstracts. Due to the similarities among different LLMs, experiments show that current classifiers struggle to accurately determine which specific model generated a given text in multi-class classification tasks. Meanwhile, variations across LLMs also result in evolving patterns of word usage in academic papers. By adopting a direct and highly interpretable linear approach and accounting for differences between models and prompts, we quantitatively assess these effects and show that real-world LLM usage is heterogeneous and dynamic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。