用语言特征差异评估大模型文本的人类相似度,发现模型表现因语域而异。
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

- 通过对比人类与模型生成文本的语言特征分布,评估人类相似性。
- 七款开源模型在五种语域中均偏离人类基准,无一接近。
- 模型接近人类程度取决于语域,而非模型大小,适合关注语言风格的研究者。
尽管事实正确性和任务表现长期是大语言模型研究的重点,但其生成文本在语言层面是否具有人类特质这一根本问题仍被忽视。从语料语言学角度看,语言生成本质上依赖语境,不同交际语境会导致语言特征频率和共现模式的差异。即使内容正确,若不符合这些模式,仍会降低人类读者的接受度。本文提出一种上下文感知的评估框架,通过双样本检验比较特定语域下人类参考语料库与对应模型生成语料库的语言特征分布。采用最大均值差异(MMD)和Biber提出的67个词法-语法特征进行实现。实验中,我们在五个涵盖不同语域的英文数据集上,对比了七款指令微调的开源模型与人类基线。结果显示,所有测试设置下,大模型均偏离人类基线;而最接近人类的语言表现取决于语域,并不随模型规模增大而提升。
原文摘要 · Abstract (English)
While factual correctness and task-performance have been in focus of Large Language Model (LLM) research for a long time, the fundamental question of how human-like generated texts are on a linguistic level has been underexplored. From a corpus-linguistic perspective, language production is inherently context-dependent, with distinct communicative contexts giving rise to differences in frequencies and co-occurrence patterns of linguistic features. A text failing to adhere to these patterns can be content-wise correct, but still be unfavorable to human readers. In this work, we propose a context-aware evaluation framework in which human-likeness is assessed using a two-sample problem between the linguistic feature distribution of a human reference corpus for a given register and a corresponding LLM-generated corpus. We implement this framework using the Maximum Mean Discrepancy (MMD) and the 67 lexico-grammatical features introduced by Biber, which are commonly applied in corpus linguistics. In our experiments, we compare seven instruction-tuned, open-source models across five English-language datasets spanning distinct registers against a human baseline. While across all tested setups, LLMs deviate from the human baseline, which models are closest to human language depends on the register and is not dictated by model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。