arXiv:2410.16107cs.CL2024-10被引 124

对比大模型与人类写作,发现风格差异仍明显

Do LLMs write like humans? Variation in grammatical and rhetorical styles

  • 用双语语料库对比人类与大模型的语法修辞特征
  • 指令微调模型风格差异更显著,大模型仍难模仿人类变体
  • 揭示可识别的深层语言模式,适合内容检测研究者

大型语言模型(LLMs)能生成符合语法、遵循指令、回答问题并解决问题的文本。随着模型能力提升,其输出与人类写作之间的区别变得越来越模糊。尽管过去研究发现了词汇选择和标点等表层特征的差异,并开发了分类器来检测模型生成内容,但尚未系统研究过大模型的修辞风格。本文使用Llama 3和GPT-4o的多个版本,从常见提示中构建了平行的人类与大模型写作语料库。基于Douglas Biber的词汇、语法和修辞特征体系,我们识别出大模型与人类之间以及不同大模型之间的系统性差异。这些差异在从小模型到大模型的演进过程中依然存在,且指令微调模型的差异大于基础模型。这一发现表明,尽管具备先进能力,大模型仍难以完全匹配人类的语言风格多样性。关注更高级的语用特征,有助于识别此前未被察觉的大模型行为模式。

原文摘要 · Abstract (English)

Large language models (LLMs) are capable of writing grammatical text that follows instructions, answers questions, and solves problems. As they have advanced, it has become difficult to distinguish their output from human-written text. While past research has found some differences in surface features such as word choice and punctuation, and developed classifiers to detect LLM output, none has studied the rhetorical styles of LLMs. Using several variants of Llama 3 and GPT-4o, we construct two parallel corpora of human- and LLM-written texts from common prompts. Using Douglas Biber's set of lexical, grammatical, and rhetorical features, we identify systematic differences between LLMs and humans and between different LLMs. These differences persist when moving from smaller models to larger ones, and are larger for instruction-tuned models than base models. This observation of differences demonstrates that despite their advanced abilities, LLMs struggle to match human stylistic variation. Attention to more advanced linguistic features can hence detect patterns in their behavior not previously recognized.

风格分析大模型检测语言学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。