评估AI生成内容的创意多样性崩溃,发现主流大模型存在过度重复问题。
Ex Ante Evaluation of AI-Induced Idea Diversity Collapse
- 将创意视为可拥挤资源,通过模型自生成对比人类基准,量化重复程度。
- 三款前沿大模型在故事、标语等任务中均出现多样性不足,超过临界值Δ。
- 提出可操作的评估方法,指导设计阶段降低内容重复,适合创意AI开发者使用。
当前创意AI主要以个体产出质量评估,但创意价值取决于群体多样性:当多人产出相似内容时,单个创意价值下降。现有评估忽视了群体层面的重复风险,导致AI可能提升个体质量却加剧集体拥挤。本文提出一种无需人机交互数据的相对评估框架,通过将创意建模为可拥挤资源,仅基于模型生成与人类基准对比,即可识别源级拥挤。引入超额拥挤系数Δ和人类相对多样性比ρ,证明ρ≥1为无过度拥挤的基准条件,并将Δ与具有暴露依赖冗余成本的采纳博弈关联。在短篇故事、营销标语及替代用途任务中,三款前沿大模型在多个拥挤核上均低于该基准。样本量合理时估计结果稳定。更重要的是,生成协议变体表明,通过针对性设计可减少拥挤,使多样性崩溃成为开发阶段可干预的评估目标。
原文摘要 · Abstract (English)
Creative AI systems are typically evaluated at the level of individual utility, yet creative outputs are consumed in populations: an idea loses value when many others produce similar ones. This creates an evaluation blind spot, as AI can improve individual outputs while increasing population-level crowding. We introduce a human-relative framework for benchmarking AI-induced human diversity collapse without requiring human-AI interaction data, providing an ex ante protocol to estimate crowding risk from model-only generations and matched unaided human baselines. By modeling ideas as congestible resources, we show that source-level crowding is identifiable from within-distribution comparisons, yielding an excess-crowding coefficient $Δ$ and a human-relative diversity ratio $ρ$. We show that $ρ\ge1$ is the no-excess-crowding parity condition and connect $Δ$ to an adoption game with exposure-dependent redundancy costs. Across short stories, marketing slogans, and alternative-uses tasks, three frontier LLMs fall below parity across crowding kernels. Estimates stabilize with feasible model-only sample sizes. Importantly, generation-protocol variants show that crowding can be reduced through targeted design, making diversity collapse an actionable, development-time evaluation target for population-aware creative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。