让故事中的人物脸型一致,支持多人多场景连贯生成
IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

- 用迭代发现与重去噪注入技术保持人物身份一致
- 在ConsiStory-Human上人脸一致性显著优于现有方法
- 适合需要长期连贯叙事或动态角色组合的创作场景
当前视觉生成模型可从文本生成具有一致角色的故事图像,但以人类为中心的故事生成面临更多挑战,如保持人物面部细节和多样性的一致性,以及跨图像协调多个角色。本文提出IdentityStory框架,实现多序列图像中人物身份的持续一致。通过驯化保留身份的生成器,框架包含两个核心组件:迭代身份发现,用于提取连贯的角色身份;重去噪身份注入,通过重去噪方式注入身份信息同时保留上下文。在ConsiStory-Human基准上的实验表明,IdentityStory在人脸一致性方面显著优于现有方法,并支持多角色组合。该框架还展现出生成无限长度故事和动态角色组合的强大潜力。
原文摘要 · Abstract (English)
Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper presents IdentityStory, a framework for human-centric story generation that ensures consistent character identity across multiple sequential images. By taming identity-preserving generators, the framework features two key components: Iterative Identity Discovery, which extracts cohesive character identities, and Re-denoising Identity Injection, which re-denoises images to inject identities while preserving desired context. Experiments on the ConsiStory-Human benchmark demonstrate that IdentityStory outperforms existing methods, particularly in face consistency, and supports multi-character combinations. The framework also shows strong potential for applications such as infinite-length story generation and dynamic character composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。