arXiv:2506.18900cs.CV2025-06被引 1

用多智能体自动修复图文生成中的角色不一致问题

Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models

  • 设计协作式多智能体框架,逐面板检测并修正视觉不一致
  • 在多面板故事生成中显著提升角色与物体的一致性表现
  • 适配多种扩散模型,无需重生成全程,适合长序列叙事生成

故事可视化已成为一项热门任务,即通过多面板图像描绘叙事场景。该任务的核心挑战在于保持视觉一致性,尤其是角色与物体在整个故事中的持续性和演变。尽管扩散模型取得进展,现有方法仍难以维持关键角色属性,导致叙事不连贯。本文提出一种协作式多智能体框架,可自主识别、修正并优化多面板故事视觉内容中的不一致。智能体在迭代循环中运行,实现细粒度的面板级更新,无需重新生成整个序列。该框架具有模型无关性,可灵活集成于多种扩散模型,包括Flux等修正流变换器及Stable Diffusion等潜在扩散模型。定量与定性实验表明,本方法在多面板一致性方面优于现有技术。

原文摘要 · Abstract (English)

Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects persist and evolve throughout the story. Despite recent advances in diffusion models, current approaches often fail to preserve key character attributes, leading to incoherent narratives. In this work, we propose a collaborative multi-agent framework that autonomously identifies, corrects, and refines inconsistencies across multi-panel story visualizations. The agents operate in an iterative loop, enabling fine-grained, panel-level updates without re-generating entire sequences. Our framework is model-agnostic and flexibly integrates with a variety of diffusion models, including rectified flow transformers such as Flux and latent diffusion models such as Stable Diffusion. Quantitative and qualitative experiments show that our method outperforms prior approaches in terms of multi-panel consistency.

故事生成扩散模型一致性修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。