arXiv:2608.05576cs.CLcs.CY2026-08

提出新框架,量化大模型生成内容的文化多样性差距。

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

论文配图:Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation
图 1 · 摘自论文原文
  • 用人类写作分布做基准,衡量大模型输出的覆盖范围。
  • 发现当前大模型内容虽合理但过于集中,缺乏多样性。
  • 适合关注生成内容丰富性与文化广度的研究者使用。

当大型语言模型(LLM)创作《哈利·波特》同人小说时,会稳定生成霍格沃茨世界中的核心元素,如标志性地点和角色;而人类创作的同人作品不仅包含这些基础内容,还融入风格各异的表达和关系多样的情节。这种大模型与人类创作之间的差异在多个领域普遍存在:大模型倾向于产生‘平均’内容,而人类写作则涵盖更广泛分布。现有研究已指出这一分布差距的存在,但尚无系统方法进行测量。本文提出一种基于人类写作的评估框架,利用特定主题下人类创作的实际分布,量化大模型生成内容的分布广度。我们引入两个指标——LLM覆盖率(LLM-Cov)和边界内率(IBR),将内容合理性与分布多样性分离开来。在创意构思与叙事生成任务中,我们发现当前大模型生成的内容虽合理但范围狭窄,集中于人类响应空间的中心区域。该框架可帮助研究人员更准确评估大模型生成内容的‘文化覆盖范围’。

原文摘要 · Abstract (English)

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce "average" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional "gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its "cultural reach".

大模型生成多样性评估文化覆盖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。