无需训练即可保持文本生成图像的主体一致性
StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
- 通过跨图像注意力共享动态对齐主体特征
- 区域特征调和提升相似细节的一致性
- 兼容预训练模型,适合创意图像生成
使用文本到图像扩散模型生成连贯的视觉故事时,保持各场景中主体的一致性是关键挑战。现有方法通常依赖微调或重训练,计算成本高且可能破坏模型原有能力。本文提出一种无需训练的方法,通过引入掩码跨图像注意力共享,动态对齐一批图像中的主体特征,并结合区域特征调和,优化视觉相似细节以增强主体一致性。实验表明,该方法在多种场景下成功生成一致的主体,同时保留了扩散模型的创作能力。
原文摘要 · Abstract (English)
Generating a coherent sequence of images that tells a visual story, using text-to-image diffusion models, often faces the critical challenge of maintaining subject consistency across all story scenes. Existing approaches, which typically rely on fine-tuning or retraining models, are computationally expensive, time-consuming, and often interfere with the model's pre-existing capabilities. In this paper, we follow a training-free approach and propose an efficient consistent-subject-generation method. This approach works seamlessly with pre-trained diffusion models by introducing masked cross-image attention sharing to dynamically align subject features across a batch of images, and Regional Feature Harmonization to refine visually similar details for improved subject consistency. Experimental results demonstrate that our approach successfully generates visually consistent subjects across a variety of scenarios while maintaining the creative abilities of the diffusion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。