用对话结构分析法量化大模型文本的重复性。
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
- 基于问题讨论框架(QUD)抽象出话语结构特征。
- 发现大模型生成文本结构重复率高于人类,即使内容不同。
- 适合关注生成文本多样性与真实性的研究者使用。
随着大语言模型在各类写作任务中能力提升,其生成内容缺乏独特性和创造性成为主要短板。尽管模型能覆盖多样话题,但文本间普遍存在重复感,我们通过相似性度量方法对其进行形式化和量化。这种熟悉感源于底层话语结构的持续存在。然而,现有依赖词汇重叠和句法模式的相似性度量主要捕捉内容相似性,难以检测结构相似性。本文引入基于问题下讨论(QUD)与问题语义的语言学理论,构建话语进展的抽象框架,并据此开发了新度量方法 QUDsim,可识别文档间的对话结构相似性。实验表明,大模型在不同样本间频繁复用话语结构(远超人类),即便内容不同;且其使用的结构类型也与人类作者显著不同。
原文摘要 · Abstract (English)
As large language models become increasingly capable at various writing tasks, their weakness at generating unique and creative content becomes a major liability. Although LLMs have the ability to generate text covering diverse topics, there is an overall sense of repetitiveness across texts that we aim to formalize and quantify via a similarity metric. The familiarity between documents arises from the persistence of underlying discourse structures. However, existing similarity metrics dependent on lexical overlap and syntactic patterns largely capture $\textit{content}$ overlap, thus making them unsuitable for detecting $\textit{structural}$ similarities. We introduce an abstraction based on linguistic theories in Questions Under Discussion (QUD) and question semantics to help quantify differences in discourse progression. We then use this framework to build $\textbf{QUDsim}$, a similarity metric that can detect discursive parallels between documents. Using QUDsim, we find that LLMs often reuse discourse structures (more so than humans) across samples, even when content differs. Furthermore, LLMs are not only repetitive and structurally uniform, but are also divergent from human authors in the types of structures they use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。