分析论文风格发现,大模型多用于统一润色而非生成内容。
GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
- 用贝叶斯分类器识别由GPT生成的文本风格
- 不同阈值下生成语言不显著影响段落风格分割
- 适合关注学术写作中AI使用方式的研究者
2022年底大型语言模型(LLMs)的兴起对学术写作带来影响,威胁可信度并引发机构不确定性。本研究旨在判断大模型是用于生成核心内容,还是仅用于编辑(如检查语法错误或不当表述)。我们通过分析arXiv论文的风格分割,采用可变PELT阈值与基于GPT生成文本训练的贝叶斯分类器进行对比。结果表明,由大模型生成的语言无法预测风格分割,说明作者在使用大模型时往往整体性地应用,从而降低了学术预印本中引入幻觉的风险。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) in late 2022 has impacted academic writing, threatening credibility, and causing institutional uncertainty. We seek to determine the degree to which LLMs are used to generate critical text as opposed to being used for editing, such as checking for grammar errors or inappropriate phrasing. In our study, we analyze arXiv papers for stylistic segmentation, which we measure by varying a PELT threshold against a Bayesian classifier trained on GPT-regenerated text. We find that LLM-attributed language is not predictive of stylistic segmentation, suggesting that when authors use LLMs, they do so uniformly, reducing the risk of hallucinations being introduced into academic preprints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。