arXiv:2510.13787cs.CV2025-10

让扩散模型讲故事时更懂上下文,自动判断何时参考旧图。

Adaptive Visual Conditioning for Semantic Consistency in Diffusion-Based Story Continuation

  • 根据当前文本动态决定是否用旧图像做参考
  • 在冲突场景下仍保持语义一致,生成质量优于基线
  • 适合需要连贯视觉叙事的创作任务

故事续写旨在生成与当前文本描述及先前图像序列保持连贯的下一帧图像。核心挑战在于有效利用先前视觉信息的同时,确保与当前文本输入的语义对齐。本文提出基于扩散模型的自适应视觉条件框架 AVC(Adaptive Visual Conditioning)。AVC 使用 CLIP 模型从历史帧中检索最语义相关的图像。当未找到足够相关的图像时,AVC 会自适应地仅在扩散过程早期限制先验视觉影响,避免引入误导性或无关信息。此外,通过大语言模型重新标注噪声数据集,提升文本监督质量与语义一致性。定量结果与人工评估表明,相比强基线,AVC 在连贯性、语义一致性和视觉保真度上表现更优,尤其在历史图像与当前输入冲突的复杂情况下。

原文摘要 · Abstract (English)

Story continuation focuses on generating the next image in a narrative sequence so that it remains coherent with both the ongoing text description and the previously observed images. A central challenge in this setting lies in utilizing prior visual context effectively, while ensuring semantic alignment with the current textual input. In this work, we introduce AVC (Adaptive Visual Conditioning), a framework for diffusion-based story continuation. AVC employs the CLIP model to retrieve the most semantically aligned image from previous frames. Crucially, when no sufficiently relevant image is found, AVC adaptively restricts the influence of prior visuals to only the early stages of the diffusion process. This enables the model to exploit visual context when beneficial, while avoiding the injection of misleading or irrelevant information. Furthermore, we improve data quality by re-captioning a noisy dataset using large language models, thereby strengthening textual supervision and semantic alignment. Quantitative results and human evaluations demonstrate that AVC achieves superior coherence, semantic consistency, and visual fidelity compared to strong baselines, particularly in challenging cases where prior visuals conflict with the current input.

扩散模型故事续写视觉一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。