通过布局条件实现角色一致性可控的连贯故事生成
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
- 利用布局信息引导帧间细粒度交互,增强故事连贯性
- 在100万张高分辨率图像上训练,支持位置/服装/表情等精确控制
- 适用于需要角色一致性与细节控制的动画/影视生成任务
近期,涉及生成一致角色的故事生成任务受到广泛关注。然而,现有方法无论是否需训练,仍因缺乏细粒度引导和帧间交互而难以保持角色一致性。此外,该领域高质量数据稀缺,导致对角色位置、外貌、服饰、表情和姿态等关键细节的精准控制困难,制约了进一步发展。本文证明,布局条件(如角色位置和详细属性)能有效促进帧间细粒度交互,不仅增强生成序列的一致性,还可精确控制角色的多种特征。基于此,我们提出全新的布局可切换故事生成任务。为解决缺乏标注布局数据的问题,我们构建了包含超过100万张720p及以上分辨率图像的Lay2Story-1M数据集,源自约11,300小时卡通视频。在此基础上,我们创建了包含3,000个提示的Lay2Story-Bench基准。同时,我们提出基于扩散变换器(DiTs)架构的Lay2Story框架,用于布局可切换故事生成。定性和定量实验表明,我们的方法优于先前的最先进(SOTA)技术,在一致性、语义相关性和美学质量方面均取得最佳表现。
原文摘要 · Abstract (English)
Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interaction. Additionally, the scarcity of high-quality data in this field makes it difficult to precisely control storytelling tasks, including the subject's position, appearance, clothing, expression, and posture, thereby hindering further advancements. In this paper, we demonstrate that layout conditions, such as the subject's position and detailed attributes, effectively facilitate fine-grained interactions between frames. This not only strengthens the consistency of the generated frame sequence but also allows for precise control over the subject's position, appearance, and other key details. Building on this, we introduce an advanced storytelling task: Layout-Togglable Storytelling, which enables precise subject control by incorporating layout conditions. To address the lack of high-quality datasets with layout annotations for this task, we develop Lay2Story-1M, which contains over 1 million 720p and higher-resolution images, processed from approximately 11,300 hours of cartoon videos. Building on Lay2Story-1M, we create Lay2Story-Bench, a benchmark with 3,000 prompts designed to evaluate the performance of different methods on this task. Furthermore, we propose Lay2Story, a robust framework based on the Diffusion Transformers (DiTs) architecture for Layout-Togglable Storytelling tasks. Through both qualitative and quantitative experiments, we find that our method outperforms the previous state-of-the-art (SOTA) techniques, achieving the best results in terms of consistency, semantic correlation, and aesthetic quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。