arXiv:2512.19012cs.CL2025-12被引 2

首个六维剧本续写评估框架,全面衡量剧情连贯性与戏剧张力。

DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation

  • 构建六维评估体系:格式规范、叙事效率、角色一致、情感深度、逻辑连贯、冲突处理
  • 在1103个剧本上评测8个顶尖模型,252次对比中有65.9%结果显著
  • 结合规则+大模型+人工验证,可为创作模型提供精准优化方向

剧本续写需保持角色一致性、推动情节连贯并维持戏剧结构,但现有基准未能全面评估这些能力。我们提出DramaBench,首个大规模六维评估框架,涵盖格式标准、叙事效率、角色一致性、情感深度、逻辑一致性和冲突处理六个独立维度。该框架融合规则分析、LLM标注与统计指标,确保评估客观可复现。我们在1,103个剧本(共8,824次评估)上对8个前沿语言模型进行全面评测,进行252次成对比较,其中65.9%具有统计显著性,并通过188个剧本的人工验证,三维度达成强一致性。消融实验表明六维度均捕捉独立质量特征(平均|r|=0.020)。DramaBench为模型改进提供可操作的维度反馈,确立了创意写作评估的新标准。

原文摘要 · Abstract (English)

Drama script continuation requires models to maintain character consistency, advance plot coherently, and preserve dramatic structurecapabilities that existing benchmarks fail to evaluate comprehensively. We present DramaBench, the first large-scale benchmark for evaluating drama script continuation across six independent dimensions: Format Standards, Narrative Efficiency, Character Consistency, Emotional Depth, Logic Consistency, and Conflict Handling. Our framework combines rulebased analysis with LLM-based labeling and statistical metrics, ensuring objective and reproducible evaluation. We conduct comprehensive evaluation of 8 state-of-the-art language models on 1,103 scripts (8,824 evaluations total), with rigorous statistical significance testing (252 pairwise comparisons, 65.9% significant) and human validation (188 scripts, substantial agreement on 3/5 dimensions). Our ablation studies confirm all six dimensions capture independent quality aspects (mean | r | = 0.020). DramaBench provides actionable, dimensionspecific feedback for model improvement and establishes a rigorous standard for creative writing evaluation.

剧本生成评估框架多维评测LLM评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。