无需训练即可生成连贯图文,适合多提示叙事场景
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- 基于尺度自回归模型,通过提示替换与注意力引导保持一致性
- 生成速度比最快现有模型快6倍以上,每图仅需1.72秒
- 无需微调,直接推理,适合真实世界视觉叙事应用
我们提出Infinite-Story,一种无需训练的连贯文本到图像(T2I)生成框架,专为多提示叙事场景设计。基于尺度自回归模型,该方法解决身份不一致和风格不一致两大挑战。提出三种互补技术:身份提示替换,缓解文本编码器中的上下文偏差,实现跨提示的身份对齐;以及统一注意力引导机制,包含自适应风格注入与同步引导适配,联合强化全局风格与身份外观一致性,同时保持提示忠实度。与需要微调或推理缓慢的扩散模型不同,Infinite-Story完全在测试阶段运行,实现多样提示下的高身份与风格一致性。大量实验表明,该方法达到当前最优生成性能,推理速度较现有最快一致T2I模型提升超6倍(每图1.72秒),凸显其有效性和实用性。
原文摘要 · Abstract (English)
We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6X faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。