arXiv:2507.11533cs.CV2025-07ICCV被引 24

让文字生成图像中的人物在多场景下保持细节一致

CharaConsist: Fine-Grained Consistent Character Generation

  • 用点追踪注意力和自适应标记合并实现前后景解耦控制
  • 支持固定场景连续镜头与跨场景离散镜头的精细一致性
  • 首个适配文本到图像DiT模型的一致性生成方法

在文本到图像生成中,保持同一角色在一系列画面中身份一致具有重要应用价值。尽管已有无需训练的方法提升生成主体的一致性,但现有方法仍存在背景细节无法保持、角色大幅运动时身份与服装细节不一致等问题。为此,我们提出CharaConsist,通过点追踪注意力与自适应标记合并,并实现前后景的解耦控制。该方法可实现前后景的细粒度一致性,支持在同一场景内连续镜头或不同场景间离散镜头中生成单一角色。此外,CharaConsist是首个专为文本到图像DiT模型设计的一致性生成方法。其细粒度一致性能力结合最新基座模型的更强容量,能生成高质量视觉输出,扩展了在真实场景中的适用范围。代码已开源。

原文摘要 · Abstract (English)

In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have explored training-free methods to enhance the consistency of generated subjects, we observe that they suffer from the following problems. First, they fail to maintain consistent background details, which limits their applicability. Furthermore, when the foreground character undergoes large motion variations, inconsistencies in identity and clothing details become evident. To address these problems, we propose CharaConsist, which employs point-tracking attention and adaptive token merge along with decoupled control of the foreground and background. CharaConsist enables fine-grained consistency for both foreground and background, supporting the generation of one character in continuous shots within a fixed scene or in discrete shots across different scenes. Moreover, CharaConsist is the first consistent generation method tailored for text-to-image DiT model. Its ability to maintain fine-grained consistency, combined with the larger capacity of latest base model, enables it to produce high-quality visual outputs, broadening its applicability to a wider range of real-world scenarios. The source code has been released at https://github.com/Murray-Wang/CharaConsist

图像生成角色一致DiT模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。