arXiv:2504.05800cs.CVcs.LG2025-04ICLR被引 7

无需训练即可保持多角色图像一致性,解决角色混淆问题。

Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling

  • 用多模态推理和区域生成预先定位角色位置
  • 提出限定注意力与合并标记层,减少角色间干扰
  • 适合需要多角色连贯视觉叙事的创作者使用

无需训练即可实现跨图像保持同一主体一致性的文本到图像生成是近期广泛关注的方向。现有方法主要依赖跨帧自注意力机制,通过让每帧中的标记关注其他帧的标记来提升一致性。然而,该方法在处理多角色时效果不佳。我们分析发现,其根本原因在于自注意力泄露——当一个角色的标记关注其他角色时,会导致角色之间混淆(如狗看起来像鸭子)。为此,我们提出StoryBooth:一种无需训练的多角色一致性增强方法。首先利用多模态思维链推理与基于区域的生成,预先定位故事中不同角色的位置;随后通过改进的扩散模型生成结果,包含两个新模块:1)受限跨帧自注意力层,减少角色间注意力泄漏;2)标记合并层,增强细节一致性。定性和定量实验表明,该方法在多角色及细粒度特征上均优于现有最优方法。

原文摘要 · Abstract (English)

Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which improves subject-consistency by allowing tokens in each frame to pay attention to tokens in other frames during self-attention computation. While useful for single subjects, we find that it struggles when scaling to multiple characters. In this work, we first analyze the reason for these limitations. Our exploration reveals that the primary-issue stems from self-attention-leakage, which is exacerbated when trying to ensure consistency across multiple-characters. This happens when tokens from one subject pay attention to other characters, causing them to appear like each other (e.g., a dog appearing like a duck). Motivated by these findings, we propose StoryBooth: a training-free approach for improving multi-character consistency. In particular, we first leverage multi-modal chain-of-thought reasoning and region-based generation to apriori localize the different subjects across the desired story outputs. The final outputs are then generated using a modified diffusion model which consists of two novel layers: 1) a bounded cross-frame self-attention layer for reducing inter-character attention leakage, and 2) token-merging layer for improving consistency of fine-grain subject details. Through both qualitative and quantitative results we find that the proposed approach surpasses prior state-of-the-art, exhibiting improved consistency across both multiple-characters and fine-grain subject details.

文本生成图像多角色一致性扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。