用社交网络分析评估大模型故事生成能力,发现其倾向构建紧密正向关系。
Evaluating LLM Story Generation through Large-scale Network Analysis of Social Structures
- 将故事转为带符号的角色关系网络,通过密度、聚类等指标量化叙事结构。
- 1200+故事对比显示,大模型生成故事中正向紧密关系占比显著高于人类创作。
- 无需人工评分,适合大规模评估大模型在复杂创作中的潜在偏差。
评估大语言模型(LLMs)在复杂任务中的创造力通常依赖人力评估,难以扩展。本文提出一种新型可扩展方法,通过分析叙事中的潜在社会结构——带符号的角色网络,来评估大模型的故事生成能力。为验证其有效性,我们基于超过1,200个故事进行了大规模对比分析,涵盖四种领先大模型(GPT-4o、GPT-4o mini、Gemini 1.5 Pro、Gemini 1.5 Flash)及一个人类写作语料库。研究基于网络密度、聚类系数和带符号边权重等属性发现,大模型生成的故事普遍表现出对紧密正向关系的强烈偏好,这一结果与以往人工评估结论一致。该方法为评估当前及未来大模型在创造性叙事中的局限性与倾向提供了有效工具。
原文摘要 · Abstract (English)
Evaluating the creative capabilities of large language models (LLMs) in complex tasks often requires human assessments that are difficult to scale. We introduce a novel, scalable methodology for evaluating LLM story generation by analyzing underlying social structures in narratives as signed character networks. To demonstrate its effectiveness, we conduct a large-scale comparative analysis using networks from over 1,200 stories, generated by four leading LLMs (GPT-4o, GPT-4o mini, Gemini 1.5 Pro, and Gemini 1.5 Flash) and a human-written corpus. Our findings, based on network properties like density, clustering, and signed edge weights, show that LLM-generated stories consistently exhibit a strong bias toward tightly-knit, positive relationships, which aligns with findings from prior research using human assessment. Our proposed approach provides a valuable tool for evaluating limitations and tendencies in the creative storytelling of current and future LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。