测试大模型能否还原文本领域的句法特征,发现生成文本偏差明显。
Domain Regeneration: How well do LLMs match syntactic properties of text domains?
- 用提示工程让模型重生成维基和新闻文本,对比句法特性
- 生成文本的句长、可读性等分布均值偏移,长尾减少
- 适合关注生成质量与真实语料差异的研究者
近期大型语言模型性能提升,很可能伴随着对训练数据分布拟合能力的增强。本文探讨的问题是:大模型在多大程度上忠实复现了文本领域的真实句法特性?我们采用语料语言学中常见的观察方法,使用一个常用的开源大模型,对两类常出现在大模型训练数据中的自由许可英文文本——维基百科和新闻文本——进行重生成。该重生成范式使我们能在相对语义受控的环境下,考察大模型是否能忠实匹配原始人类文本领域的句法特征。研究涵盖从简单属性(如句子长度、文章可读性)到复杂高阶属性(如依存标签分布、解析深度与解析复杂度)的多层次句法抽象。结果表明,大多数生成文本的分布相较于原有人类文本,呈现均值偏移、标准差降低及长尾缩减现象。
原文摘要 · Abstract (English)
Recent improvement in large language model performance have, in all likelihood, been accompanied by improvement in how well they can approximate the distribution of their training data. In this work, we explore the following question: which properties of text domains do LLMs faithfully approximate, and how well do they do so? Applying observational approaches familiar from corpus linguistics, we prompt a commonly used, opensource LLM to regenerate text from two domains of permissively licensed English text which are often contained in LLM training data -- Wikipedia and news text. This regeneration paradigm allows us to investigate whether LLMs can faithfully match the original human text domains in a fairly semantically-controlled setting. We investigate varying levels of syntactic abstraction, from more simple properties like sentence length, and article readability, to more complex and higher order properties such as dependency tag distribution, parse depth, and parse complexity. We find that the majority of the regenerated distributions show a shifted mean, a lower standard deviation, and a reduction of the long tail, as compared to the human originals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。