用大模型生成连贯图文叙事,评估更贴近人类判断。
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
- 基于Transformer和多模态大模型,生成与图像匹配的连贯故事。
- 在VIST数据集上,新评测指标比传统方法更准确反映叙事质量。
- 适合关注图文生成、叙事理解的研究者与应用开发者。
视觉叙事是计算机视觉与自然语言处理的交叉领域,旨在从图像序列中生成连贯的叙事。本文提出VIST-GPT,一种利用最新多模态模型进展的新型方法,基于大规模视觉叙事(VIST)数据集,生成具有视觉依据且语境恰当的叙事内容。针对传统评估指标(如BLEU、METEOR、ROUGE、CIDEr)不适用于该任务的问题,我们引入两种新颖的无参考评估指标——RoViST与GROOVIST,重点衡量视觉一致性、连贯性与非冗余性。这些指标能更细致地评估叙事质量,与人工评价高度一致。
原文摘要 · Abstract (English)
Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in multimodal models, specifically adapting transformer-based architectures and large multimodal models, for the visual storytelling task. Leveraging the large-scale Visual Storytelling (VIST) dataset, our VIST-GPT model produces visually grounded, contextually appropriate narratives. We address the limitations of traditional evaluation metrics, such as BLEU, METEOR, ROUGE, and CIDEr, which are not suitable for this task. Instead, we utilize RoViST and GROOVIST, novel reference-free metrics designed to assess visual storytelling, focusing on visual grounding, coherence, and non-redundancy. These metrics provide a more nuanced evaluation of narrative quality, aligning closely with human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。