为评估和提升大模型写剧本能力,提出CML-Bench评测框架。
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
- 基于真实电影剧本构建三维度评测体系:对白连贯、角色一致、情节合理。
- 新提示策略使大模型生成剧本更符合电影逻辑,与人工评价高度一致。
- 适合影视创作、AI编剧研究者使用,推动大模型叙事能力提升。
大语言模型在生成结构化文本方面表现卓越,但难以捕捉电影剧本所需的细腻叙事与情感深度。为此,我们构建了CML-Dataset,包含(摘要,内容)对的电影标记语言数据集,其中‘内容’来自优质电影剧本片段,‘摘要’为内容简述。通过分析真实剧本中的多段连续性与叙事结构,我们识别出三个关键质量维度:对白连贯性(DC)、角色一致性(CC)和情节合理性(PR)。据此提出CML-Bench,涵盖上述维度的量化指标。该基准能有效区分高质量人工剧本与大模型生成剧本,精准定位缺陷。为进一步验证,引入CML-Instruction提示策略,通过细化角色对话与事件逻辑指令,引导大模型生成更符合电影语境的剧本。大量实验表明,该策略显著提升生成质量,结果与人工偏好高度一致。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating highly structured texts. However, while exhibiting a high degree of structural organization, movie scripts demand an additional layer of nuanced storytelling and emotional depth-the 'soul' of compelling cinema-that LLMs often fail to capture. To investigate this deficiency, we first curated CML-Dataset, a dataset comprising (summary, content) pairs for Cinematic Markup Language (CML), where 'content' consists of segments from esteemed, high-quality movie scripts and 'summary' is a concise description of the content. Through an in-depth analysis of the intrinsic multi-shot continuity and narrative structures within these authentic scripts, we identified three pivotal dimensions for quality assessment: Dialogue Coherence (DC), Character Consistency (CC), and Plot Reasonableness (PR). Informed by these findings, we propose the CML-Bench, featuring quantitative metrics across these dimensions. CML-Bench effectively assigns high scores to well-crafted, human-written scripts while concurrently pinpointing the weaknesses in screenplays generated by LLMs. To further validate our benchmark, we introduce CML-Instruction, a prompting strategy with detailed instructions on character dialogue and event logic, to guide LLMs to generate more structured and cinematically sound scripts. Extensive experiments validate the effectiveness of our benchmark and demonstrate that LLMs guided by CML-Instruction generate higher-quality screenplays, with results aligned with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。