MangaFlow实现可控制的漫画全流程生成,支持精细布局与跨面板一致性。
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

- 将漫画生成分解为规划、定位、布局、渲染等步骤,显式控制中间变量。
- 在长篇漫画中保持角色、场景一致性,跨面板一致率提升显著。
- 适合需要精细编辑布局和文字的创作者,也适用于交互式漫画设计。
端到端漫画生成是一项结构化视觉叙事任务,需完成故事拆解、角色与场景定位、页面布局设计、分镜渲染、页面组合及文字排版。现有生成模型多直接合成整页图像,将多个因素混杂于单一输出中,难以精确控制布局几何、视觉参考和跨分镜一致性。为此,我们提出MangaFlow——一个面向可控长篇漫画生成的智能体框架,将漫画创作分解为规划、定位、布局构建、参考条件渲染、组合与文本放置等阶段。通过将布局和视觉参考作为显式中间变量,MangaFlow既支持简单文本到漫画生成,也支持用户精细化控制。该设计使布局、视觉素材和文字可作为可编辑中间控件,用于调整分镜结构、参照物和文本位置。为保障长篇内容一致性,MangaFlow引入故事段落记忆机制,将段落描述与对应的角色、场景、物体参考关联,实现跨分镜复用。我们还构建了一个元基准评估布局可控性、视觉一致性与生成质量。实验表明,MangaFlow在布局遵循度和跨面板一致性上优于直接生成基线,同时支持灵活的人机协同控制。
原文摘要 · Abstract (English)
End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page layout design, panel rendering, page composition, and lettering. However, existing generative models often perform direct page synthesis, entangling these factors in a single visual output and limiting precise control over layout geometry, visual references, and cross-panel consistency. To address these limitations, we propose MangaFlow, an agentic framework for controllable long-form manga generation that decomposes manga creation into planning, grounding, layout construction, reference-conditioned rendering, composition, and text placement. By treating layout and visual references as explicit intermediate variables, MangaFlow enables both simple text-to-manga generation and more precise user-controlled manga creation. This design exposes layout, visual assets, and lettering as editable intermediate controls for refining panel geometry, references, and text placement. To support long-form consistency, MangaFlow introduces a story section memory that links section descriptions with corresponding character, scene, and object references for reuse across panels. We further present a meta-benchmark for evaluating layout controllability, visual consistency, and generation quality. Experiments show that MangaFlow improves layout adherence and cross-panel consistency over direct generation baselines while supporting flexible human control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。