分析11个大模型在8类文体中的写作风格差异,揭示模型和文体比提示词影响更大。
Interpretable Stylistic Variation in Human and LLM Writing Across Genres, Models, and Decoding Strategies

- 用语法学特征对比人类与大模型文本风格差异
- 模型类型影响大于生成策略,文体影响超过来源
- 聊天版模型风格趋同,适合研究生成文本可识别性
大型语言模型(LLMs)已能生成高度流畅、接近人类的文本,但其广泛应用也带来虚假信息、钓鱼攻击及学术滥用等风险。尽管检测机器生成文本的研究较多,但对人类与机器文本风格差异的理解仍有限。本文对11个大模型在8种不同文体和4种解码策略下生成的文本进行了大规模分析,采用Douglas Biber的词汇语法与功能特征集。结果表明:关键的语言学差异在不同生成条件下保持稳定;文体对风格的影响强于文本来源;聊天类模型在风格空间中趋于聚集;模型本身对风格的影响大于解码策略,少数例外存在。这些发现强调了模型和文体在塑造机器生成文本风格中的主导作用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are now capable of generating highly fluent, human-like text. They enable many applications, but also raise concerns such as large scale spam, phishing, or academic misuse. While much work has focused on detecting LLM-generated text, only limited work has gone into understanding the stylistic differences between human-written and machine-generated text. In this work, we perform a large scale analysis of stylistic variation across human-written text and outputs from 11 LLMs spanning 8 different genres and 4 decoding strategies using Douglas Biber's set of lexicogrammatical and functional features. Our findings reveal insights that can guide intentional LLM usage. First, key linguistic differentiators of LLM-generated text seem robust to generation conditions (e.g., prompt settings to nudge them to generate human-like text, or availability of human-written text to continue the style); second, genre exerts a stronger influence on stylistic features than the source itself; third, chat variants of the models generally appear to be clustered together in stylistic space, and finally, model has a larger effect on the style than decoding strategy, with some exceptions. These results highlight the relative importance of model and genre over prompting and decoding strategies in shaping the stylistic behavior of machine-generated text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。