构建视觉叙事基准,提升图文生成一致性与忠实度。
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
- 标注常识与话语约束,为叙事生成提供结构化知识
- 新评测指标有效提升生成结果的连贯性与对齐度
- 适合研究多图生成、视觉叙事与对齐评估的学者
视觉叙事生成将文本叙事转化为一系列图像以展现内容。然而,由于缺乏用于规划故事的知识约束,生成既忠实于输入文本又在图像间保持自洽的视觉叙事仍是一个开放挑战。本文提出新基准VinaBench,对视觉叙事样本中的常识和话语约束进行标注,为学习隐含的视觉叙事策略提供系统性框架。基于内置的叙事约束,我们进一步设计了新型评估指标,以更精准衡量生成图像的一致性及与原文本的对齐程度。在三个生成视觉模型上的实验表明,利用VinaBench的知识约束可显著提升生成视觉叙事的忠实度与连贯性。
原文摘要 · Abstract (English)
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge constraints used for planning the stories. In this work, we propose a new benchmark, VinaBench, to address this challenge. Our benchmark annotates the underlying commonsense and discourse constraints in visual narrative samples, offering systematic scaffolds for learning the implicit strategies of visual storytelling. Based on the incorporated narrative constraints, we further propose novel metrics to closely evaluate the consistency of generated narrative images and the alignment of generations with the input textual narrative. Our results across three generative vision models demonstrate that learning with VinaBench's knowledge constraints effectively improves the faithfulness and cohesion of generated visual narratives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。