让生成的故事角色始终一致且有指代关系,避免角色混乱
Generating Visual Stories with Grounded and Coreferent Characters
- 基于新构建的数据集,训练角色连贯的视觉故事生成模型
- 生成故事中角色出现更频繁且指代一致,优于现有系统
- 适合需要角色连贯性的多模态叙事任务
角色在叙事中至关重要,推动情节发展、建立情感联结并体现主题。现有视觉叙事方法更关注事件与情节,缺乏围绕特定角色的构建,导致生成故事角色缺失、模糊或错误。为此,我们提出角色中心的视觉故事生成新任务,并构建首个能生成一致且有指代关系角色的模型。该模型在基于广泛使用的VIST基准构建的新数据集上微调,通过自动化流程为VIST补充视觉与文本层面的角色共指链。同时提出新评估指标,衡量故事中角色与共指的丰富度。实验表明,相较于基线与先进系统,本模型生成的故事具有更高频率的角色重复、更强的一致性与共指性。
原文摘要 · Abstract (English)
Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story's themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative around specific characters. As a result, the generated stories feel generic, with character mentions being absent, vague, or incorrect. To mitigate these issues, we introduce the new task of character-centric story generation and present the first model capable of predicting visual stories with consistently grounded and coreferent character mentions. Our model is finetuned on a new dataset which we build on top of the widely used VIST benchmark. Specifically, we develop an automated pipeline to enrich VIST with visual and textual character coreference chains. We also propose new evaluation metrics to measure the richness of characters and coreference in stories. Experimental results show that our model generates stories with recurring characters which are consistent and coreferent to larger extent compared to baselines and state-of-the-art systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。