arXiv:2605.06030cs.CL2026-05ACL

对比两代大模型与新闻文本的语法和词汇多样性,发现新模型更规整但表达更单一。

More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs

论文配图:More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs
图 1 · 摘自论文原文
  • 用语法规则框架分析模型生成文本的句法结构与词汇类型分布。
  • 新模型相比旧模型在句法和词汇多样性上均有下降,降幅达15%以上。
  • 适合关注模型可控性与表达丰富性权衡的研究者阅读。

本研究比较了大语言模型生成文本与人类撰写的英文新闻文本,聚焦句法特性分析。基于头驱动短语结构语法(HPSG)形式化框架,我们分析了两个时代的人类新闻数据集(来自不同年份的《纽约时报》)与两代大模型生成文本的句法结构及词汇类型的分布。采用生态学与信息论中的多样性度量方法,量化了语法构造与词汇类型的变异程度。结果显示,在考察时间段内,人类新闻文本的句法特征变化微小;而较新的大模型相较于早期非指令微调模型,表现出显著更低的句法与词汇多样性。这一发现表明,尽管指令微调提升了输出连贯性与提示遵循度,却可能压缩了模型表达范围,提示未来需深入研究其对语言多样性的潜在影响。

原文摘要 · Abstract (English)

This study contributes to a growing line of research in comparing LLM-generated texts with human-authored text, in this case, English news text. We focus in particular on the evaluation of syntactic properties through formal grammar frameworks. Our analysis compares two generations of LLMs in the context of two human-authored English news datasets from two different years. Employing the Head-Driven Phrase Structure Grammar (HPSG) formalism, we investigate the distributions of syntactic structures and lexical types of AI-generated texts and contrast them with the corresponding distributions in the human-authored New York Times (NYT) articles. We use diversity metrics from ecology and information theory to quantify variation in grammatical constructions and lexical types. We show that English news text has changed little in the given time frame, while newer LLMs display reduced syntactic and, especially, lexical diversity compared to older, non-instruction-tuned models. These findings point to future work in studying effects of instruction tuning, which, while enhancing coherence and adherence to prompts, may narrow the expressive range of model output.

语言模型语法分析多样性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。