让故事生成更连贯,自动规划场景并保持一致性
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
- 用视觉语言模型引导全局到局部的场景规划
- 通过长期场景共享注意力维持多故事间场景一致性和角色多样性
- 无需训练,适合艺术、影视、游戏等创意领域
近期文本到图像模型虽已革新图像生成,但仍难以保持生成图像间概念的一致性。现有工作多关注角色一致性,却忽视了场景在叙事中的关键作用,限制了实际创造力。本文提出面向场景的故事生成,解决两大挑战:(i) 场景规划——当前方法仅依赖文本描述,无法保证场景层面的叙事连贯性;(ii) 场景一致性——跨多个故事中维持场景一致性的研究仍属空白。我们提出SceneDecorator,一种无需训练的框架,采用VLM-Guided Scene Planning实现从全局到局部的叙事连贯性,以及Long-Term Scene-Sharing Attention,以维持长期场景一致性和主体多样性。大量实验表明SceneDecorator表现优越,展现出在艺术、影视、游戏等领域释放创造力的巨大潜力。
原文摘要 · Abstract (English)
Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。