用分离注意力提升角色一致性,无需训练即可生成新角色故事图
Object Isolated Attention for Consistent Story Visualization
- 分离自注意力与交叉注意力,聚焦角色关键特征
- 无需训练,可连续生成新角色和故事
- 在角色一致性和场景合理性上优于现有方法
开放式故事可视化是一项挑战性任务,需从给定剧情生成连贯的图像序列。主要难点在于保持角色一致性的同时生成自然且符合上下文的场景,现有方法在此方面表现不佳。本文提出一种增强型Transformer模块,通过独立的自注意力与交叉注意力机制,利用预训练扩散模型的先验知识,确保逻辑合理的场景生成。隔离自注意力机制通过优化注意力图,减少对无关区域的关注,突出同一角色的关键特征;隔离交叉注意力机制独立处理每个角色特征,避免特征融合,进一步强化一致性。值得注意的是,该方法无需训练,可直接生成新角色和新故事线而无需重新调参。定性与定量评估均表明,本方法显著优于现有技术,验证了其有效性。
原文摘要 · Abstract (English)
Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character consistency while creating natural and contextually fitting scenes--an area where many existing methods struggle. In this paper, we propose an enhanced Transformer module that uses separate self attention and cross attention mechanisms, leveraging prior knowledge from pre-trained diffusion models to ensure logical scene creation. The isolated self attention mechanism improves character consistency by refining attention maps to reduce focus on irrelevant areas and highlight key features of the same character. Meanwhile, the isolated cross attention mechanism independently processes each character's features, avoiding feature fusion and further strengthening consistency. Notably, our method is training-free, allowing the continuous generation of new characters and storylines without re-tuning. Both qualitative and quantitative evaluations show that our approach outperforms current methods, demonstrating its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。