构建首个全面的故事可视化评测基准,涵盖多角色、多场景叙事
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
- 基于文学影视神话故事生成带角色引用的多镜头脚本
- 提出五维自动化指标,量化角色一致性与画面质量
- 适合研究视觉叙事、多角色生成的学者与开发者
故事可视化旨在生成连贯图像序列,忠实呈现叙事内容并匹配指定角色。现有评测基准范围狭窄,常局限于简短提示、缺乏角色引用或仅支持单图生成,难以反映真实叙事复杂性,掩盖模型真实性能。我们提出ViStoryBench,一个综合性评测基准,用于评估故事可视化模型在不同叙事结构、视觉风格和角色设定下的表现。该基准包含从文学、电影和民间故事中精选的丰富标注多镜头脚本,利用大语言模型辅助生成,并经人工验证确保连贯性和忠实度。角色引用经过精心设计,以保证跨艺术风格的一致性。提出一套多维度自动化指标,评估角色一致性、风格相似性、提示对齐度、美学质量及复制粘贴等伪影行为。这些指标通过人类评估验证,并用于测评广泛开源与商业模型,实现系统性分析,推动视觉叙事技术发展。
原文摘要 · Abstract (English)
Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to short prompts, lacking character references, or single-image cases, failing to reflect real-world narrative complexity and obscuring true model performance.We introduce ViStoryBench, a comprehensive benchmark designed to evaluate story visualization models across varied narrative structures, visual styles, and character settings. It features richly annotated multi-shot scripts derived from curated stories spanning literature, film, and folklore. Large language models assist in story summarization and script generation, with all outputs verified by humans for coherence and fidelity. Character references are carefully curated to maintain consistency across different artistic styles. ViStoryBench proposes a suite of multi-dimensional automated metrics to evaluate character consistency, style similarity, prompt alignment, aesthetic quality, and artifacts like copy-paste behavior. These metrics are validated through human studies and used to assess a broad range of open-source and commercial models, enabling systematic analysis and encouraging advances in visual storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。