用4000部网络小说评估大模型的长篇叙事能力
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
- 基于4000部中文网文,设计摘要生成故事的评测任务
- 从8个维度自动评分,区分人类佳作、热门网文与模型生成内容
- 适合研究生成式AI叙事能力或需客观评测的团队使用
评估大语言模型(LLMs)在长篇叙事方面的能力仍面临挑战,现有基准常缺乏规模、多样性或客观指标。为此,我们提出WebNovelBench,一个专为长篇小说生成设计的新基准。该基准基于超过4000部中文网络小说,将评测任务定义为从摘要生成完整故事。我们构建了包含八个叙事质量维度的多维度评估框架,通过LLM作为评判者实现自动化评分,并利用主成分分析整合得分,映射至与人类作品对比的百分位排名。实验表明,WebNovelBench能有效区分人类创作的杰作、流行网文和模型生成内容。我们对24个前沿大模型进行了全面分析,对其叙事能力进行排序,并为未来研究提供洞见。该基准提供了可扩展、可复现且数据驱动的评估方法,推动基于大模型的叙事生成发展。
原文摘要 · Abstract (English)
Robustly evaluating the long-form storytelling capabilities of Large Language Models (LLMs) remains a significant challenge, as existing benchmarks often lack the necessary scale, diversity, or objective measures. To address this, we introduce WebNovelBench, a novel benchmark specifically designed for evaluating long-form novel generation. WebNovelBench leverages a large-scale dataset of over 4,000 Chinese web novels, framing evaluation as a synopsis-to-story generation task. We propose a multi-faceted framework encompassing eight narrative quality dimensions, assessed automatically via an LLM-as-Judge approach. Scores are aggregated using Principal Component Analysis and mapped to a percentile rank against human-authored works. Our experiments demonstrate that WebNovelBench effectively differentiates between human-written masterpieces, popular web novels, and LLM-generated content. We provide a comprehensive analysis of 24 state-of-the-art LLMs, ranking their storytelling abilities and offering insights for future development. This benchmark provides a scalable, replicable, and data-driven methodology for assessing and advancing LLM-driven narrative generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。