arXiv:2508.14045cs.CLcs.CV2025-08

将图像描述与视觉叙事统一框架,提升故事连贯性。

From Image Captioning to Visual Storytelling

  • 先用图像描述模型生成每张图的文本,再用语言模型整合成连贯故事。
  • 新框架在多个评测中表现更优,训练速度更快且可复现。
  • 提出新评估指标'ideality',模拟人类叙事水平。

视觉叙事是视觉与语言间的挑战性多模态任务,目标是为一系列图像生成连贯的故事。其难点在于故事既要基于图像序列,又要具备叙事性和逻辑性。本文将视觉叙事视为图像描述的超集,提出新方法:首先使用视觉-语言模型生成输入图像的描述,再通过语言-语言方法将这些描述转化为连贯叙述。多种评估表明,统一框架显著提升了故事质量。相比以往研究,该方法训练更快,易于复现和重用。最后,本文提出新度量工具「ideality」,用于模拟结果距离理想模型的距离,实现对视觉叙事人类相似度的量化评估。

原文摘要 · Abstract (English)

Visual Storytelling is a challenging multimodal task between Vision & Language, where the purpose is to generate a story for a stream of images. Its difficulty lies on the fact that the story should be both grounded to the image sequence but also narrative and coherent. The aim of this work is to balance between these aspects, by treating Visual Storytelling as a superset of Image Captioning, an approach quite different compared to most of prior relevant studies. This means that we firstly employ a vision-to-language model for obtaining captions of the input images, and then, these captions are transformed into coherent narratives using language-to-language methods. Our multifarious evaluation shows that integrating captioning and storytelling under a unified framework, has a positive impact on the quality of the produced stories. In addition, compared to numerous previous studies, this approach accelerates training time and makes our framework readily reusable and reproducible by anyone interested. Lastly, we propose a new metric/tool, named ideality, that can be used to simulate how far some results are from an oracle model, and we apply it to emulate human-likeness in visual storytelling.

视觉叙事多模态语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。