arXiv:2506.12198cs.CV2025-06中稿 · WACV 2026被引 6

用多模态适配器让图像生成更连贯,适合讲故事的场景。

ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

  • 通过融合历史图文对提取关键上下文特征
  • 生成序列图像保持帧间一致性且贴合文本描述
  • 适合需要连贯视觉叙事的应用,如动画或故事生成

文本到图像的扩散模型已取得显著进展,但生成连贯的图像序列以实现视觉叙事仍具挑战。核心难点在于有效利用所有先前的图文对(即历史图文对),这些信息有助于维持帧间的一致性。现有自回归方法依赖全部历史图文对进行条件生成,但需大量训练;而无需训练的特定主体方法虽能保证一致性,却难以适应不同叙事提示。为此,我们提出多模态历史适配器模型ViSTA,包含(1)多模态历史融合模块,用于提取相关历史特征;(2)历史适配器,将提取特征用于条件生成。此外,在推理阶段引入显著历史选择策略,仅选取最相关的前序图文对,提升条件质量。我们还提出基于视觉问答的评估指标TIFA,更精准地衡量叙事中图文对齐程度。在StorySalon和FlintStonesSV数据集上的实验表明,ViSTA不仅能保持多帧间的一致性,还能准确匹配叙事文本描述。

原文摘要 · Abstract (English)

Text-to-image diffusion models have achieved remarkable success, yet generating coherent image sequences for visual storytelling remains challenging. A key challenge is effectively leveraging all previous text-image pairs, referred to as history text-image pairs, which provide contextual information for maintaining consistency across frames. Existing auto-regressive methods condition on all past image-text pairs but require extensive training, while training-free subject-specific approaches ensure consistency but lack adaptability to narrative prompts. To address these limitations, we propose a multi-modal history adapter for text-to-image diffusion models, \textbf{ViSTA}. It consists of (1) a multi-modal history fusion module to extract relevant history features and (2) a history adapter to condition the generation on the extracted relevant features. We also introduce a salient history selection strategy during inference, where the most salient history text-image pair is selected, improving the quality of the conditioning. Furthermore, we propose to employ a Visual Question Answering-based metric TIFA to assess text-image alignment in visual storytelling, providing a more targeted and interpretable assessment of generated images. Evaluated on the StorySalon and FlintStonesSV dataset, our proposed ViSTA model is not only consistent across different frames, but also well-aligned with the narrative text descriptions.

视觉叙事图像生成多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。