人类与大模型在学术写作中相互适应,导致文本检测更难。
Human-LLM Coevolution: Evidence from Academic Writing
- 通过分析论文摘要词频变化,发现作者已调整用词以规避AI痕迹。
- 如' delved '等曾被指过度使用的词频率下降,但' significant '等仍持续上升。
- 研究提示需关注高频词演变,对检测真实机器生成内容有启示。
通过对arXiv论文摘要的统计分析,我们发现自2024年初被指出后,一些此前被认定为ChatGPT过度使用的词汇(如' delved ')频率显著下降;而另一些受ChatGPT偏好的词汇(如' significant ')频率反而持续上升。这些现象表明,部分学术作者已开始调整其使用大语言模型(LLMs)的方式,例如选择或修改模型输出内容。这种人与大模型的共进化与协作关系,使得在真实场景中检测机器生成文本面临新的挑战。仅通过词频变化评估大模型对学术写作的影响仍具可行性,且应更加关注那些本就高频、甚至因大模型偏好而频率下降的词汇。
原文摘要 · Abstract (English)
With a statistical analysis of arXiv paper abstracts, we report a marked drop in the frequency of several words previously identified as overused by ChatGPT, such as "delve", starting soon after they were pointed out in early 2024. The frequency of certain other words favored by ChatGPT, such as "significant", has instead kept increasing. These phenomena suggest that some authors of academic papers have adapted their use of large language models (LLMs), for example, by selecting outputs or applying modifications to the LLM-generated content. Such coevolution and cooperation of humans and LLMs thus introduce additional challenges to the detection of machine-generated text in real-world scenarios. Estimating the impact of LLMs on academic writing by examining word frequency remains feasible, and more attention should be paid to words that were already frequently employed, including those that have decreased in frequency due to LLMs' disfavor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。