arXiv:2509.04446cs.CV2025-09AAAI被引 1

零样本生成连贯故事图,支持精细编辑与一致性控制。

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models

  • 基于扩散模型实现无需训练的故事图像生成与编辑。
  • 可对多帧图像进行细粒度修改且保持叙事一致性。
  • 适合需要快速创作和迭代视觉故事的创作者。

文本到图像的扩散模型在多个领域展现出生成多样化、细节丰富图像的强大能力,故事可视化正成为一项极具前景的应用。然而,随着其在真实创意场景中的应用增多,如何提供增强的控制力、精细化调整能力以及生成后一致性的图像修改变得尤为关键。现有方法往往难以在保持多帧间视觉与叙事一致性的同时实现粗粒度或细粒度的编辑,阻碍了创作者无缝构建和优化视觉故事。为此,我们提出 Plot'n Polish,一个零样本框架,能够实现一致性的故事生成,并在不同细节层级上提供对故事视觉表达的精细控制。

原文摘要 · Abstract (English)

Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in real-world creative domains increases, the need for providing enhanced control, refinement, and the ability to modify images post-generation in a consistent manner becomes an important challenge. Existing methods often lack the flexibility to apply fine or coarse edits while maintaining visual and narrative consistency across multiple frames, preventing creators from seamlessly crafting and refining their visual stories. To address these challenges, we introduce Plot'n Polish, a zero-shot framework that enables consistent story generation and provides fine-grained control over story visualizations at various levels of detail.

故事生成扩散模型图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。