让故事中角色保持一致,还能灵活换姿势、背景更生动。
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection

- 用帧间同步与打乱引用策略,让角色形象跨画面保持一致。
- 在10万张图像的数据集上训练,生成故事连贯性更强。
- 适合需要角色一致性高的视频/漫画生成场景。
近期文本到图像扩散模型的发展推动了连续故事图像生成的研究。本文提出Storynizor,一种能生成连贯故事、保持角色高度一致、有效分离前景与背景并支持多样化姿态变化的模型。其核心创新在于两个模块:ID-Synchronizer通过自掩码注意力机制和帧间掩码感知损失提升角色生成的一致性,清晰呈现姿态与背景;ID-Injector采用打乱参考策略(SRS)将身份特征注入特定位置,增强基于身份的角色一致性生成。此外,为支持模型训练,我们构建了一个新数据集StoryDB,包含100,000张图像,涵盖单人与多人在不同环境、布局和动作下的详细描述。实验表明,相比其他特定角色方法,Storynizor在生成连贯故事方面表现更优,具备高保真角色一致性、灵活姿态和生动背景。
原文摘要 · Abstract (English)
Recent advances in text-to-image diffusion models have spurred significant interest in continuous story image generation. In this paper, we introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency, effective foreground-background separation, and diverse pose variation. The core innovation of Storynizor lies in its key modules: ID-Synchronizer and ID-Injector. The ID-Synchronizer employs an auto-mask self-attention module and a mask perceptual loss across inter-frame images to improve the consistency of character generation, vividly representing their postures and backgrounds. The ID-Injector utilize a Shuffling Reference Strategy (SRS) to integrate ID features into specific locations, enhancing ID-based consistent character generation. Additionally, to facilitate the training of Storynizor, we have curated a novel dataset called StoryDB comprising 100, 000 images. This dataset contains single and multiple-character sets in diverse environments, layouts, and gestures with detailed descriptions. Experimental results indicate that Storynizor demonstrates superior coherent story generation with high-fidelity character consistency, flexible postures, and vivid backgrounds compared to other character-specific methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。