无需训练,通过迭代优化实现百帧长故事的连贯可视化
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
- 用全局参考注意力模块,逐轮融合所有参考图提升一致性
- 在100帧长故事上实现最优语义连贯性和细节交互
- 适合需要高质量长序列生成的视觉叙事任务
本文提出Story-Iter,一种无需训练的迭代范式,用于增强长故事生成。与依赖固定参考图的方法不同,该方法通过外部迭代机制,在每轮中利用前一轮的所有参考图持续优化生成图像。为此,我们设计了一个即插即用、无需训练的全局参考交叉注意力(GRCA)模块,通过全局嵌入建模所有参考帧,确保长序列的语义一致性。通过逐步融入整体视觉上下文和文本约束,该迭代范式实现了细粒度交互的精准生成,逐步优化故事可视化效果。在官方故事可视化数据集及自建的长故事基准上的大量实验表明,Story-Iter在长达100帧的长故事生成中,性能达到当前最佳,显著提升了语义一致性和细粒度交互能力。
原文摘要 · Abstract (English)
This paper introduces Story-Iter, a new training-free iterative paradigm to enhance long-story generation. Unlike existing methods that rely on fixed reference images to construct a complete story, our approach features a novel external iterative paradigm, extending beyond the internal iterative denoising steps of diffusion models, to continuously refine each generated image by incorporating all reference images from the previous round. To achieve this, we propose a plug-and-play, training-free global reference cross-attention (GRCA) module, modeling all reference frames with global embeddings, ensuring semantic consistency in long sequences. By progressively incorporating holistic visual context and text constraints, our iterative paradigm enables precise generation with fine-grained interactions, optimizing the story visualization step-by-step. Extensive experiments in the official story visualization dataset and our long story benchmark demonstrate that Story-Iter's state-of-the-art performance in long-story visualization (up to 100 frames) excels in both semantic consistency and fine-grained interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。