arXiv:2507.17853cs.CVcs.AI2025-07被引 5

无需训练即可提升文本生成图像的细节一致性。

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

  • 分阶段生成:将复杂提示拆解为子提示逐步构建图像。
  • 在T2I-CompBench上优于现有方法,多对象场景提升显著。
  • 适合需要精准属性绑定的复杂图像生成任务。

文本到图像(T2I)生成虽取得显著进展,但在处理含多个主体及其独立属性的复杂提示时仍面临挑战。受人类绘画过程启发——先构图再添细节——我们提出Detail++,一种无需训练的框架,引入新型渐进式细节注入(PDI)策略。具体地,将复杂提示分解为一系列简化子提示,分阶段引导生成过程。该分步生成利用自注意力机制的布局控制能力,先确保全局结构,再进行精细优化。为实现属性与对应主体的准确绑定,我们利用交叉注意力机制,并在推理时引入质心对齐损失,减少绑定噪声,提升属性一致性。在T2I-CompBench及新构建的风格组合基准上的实验表明,Detail++显著优于现有方法,尤其在多对象与复杂风格条件下表现突出。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct attributes. Inspired by the human drawing process, which first outlines the composition and then incrementally adds details, we propose Detail++, a training-free framework that introduces a novel Progressive Detail Injection (PDI) strategy to address this limitation. Specifically, we decompose a complex prompt into a sequence of simplified sub-prompts, guiding the generation process in stages. This staged generation leverages the inherent layout-controlling capacity of self-attention to first ensure global composition, followed by precise refinement. To achieve accurate binding between attributes and corresponding subjects, we exploit cross-attention mechanisms and further introduce a Centroid Alignment Loss at test time to reduce binding noise and enhance attribute consistency. Extensive experiments on T2I-CompBench and a newly constructed style composition benchmark demonstrate that Detail++ significantly outperforms existing methods, particularly in scenarios involving multiple objects and complex stylistic conditions.

图像生成扩散模型细节增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。