通过动态重加权注意力图,抑制文本到图像模型的语义泄露。
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- 推理时直接干预注意力机制,无需优化或外部输入。
- 在1130个样本上验证,显著降低语义泄露,保持生成质量。
- 适合关注生成可控性与语义精确性的研究人员。
文本到图像(T2I)模型快速发展,但仍存在语义泄露问题,即不同实体间语义相关特征的无意传递。现有方法多依赖优化或外部输入。本文提出DeLeaker,一种轻量级、无需优化的推理时方法,通过动态重加权注意力图,抑制跨实体过度交互,强化各实体身份。为系统评估,我们构建了首个专注于语义泄露的SLIM数据集,包含1,130个经人工验证的样本,覆盖多种场景,并设计新的自动评估框架。实验表明,DeLeaker在所有基线(包括使用外部信息的)之上持续表现更优,有效缓解泄漏,同时不损害生成保真度与质量。结果凸显注意力控制的价值,为更语义精确的T2I模型铺路。
原文摘要 · Abstract (English)
Text-to-Image (T2I) models have advanced rapidly, yet they remain vulnerable to semantic leakage, the unintended transfer of semantically related features between distinct entities. Existing mitigation strategies are often optimization-based or dependent on external inputs. We introduce DeLeaker, a lightweight, optimization-free inference-time approach that mitigates leakage by directly intervening on the model's attention maps. Throughout the diffusion process, DeLeaker dynamically reweights attention maps to suppress excessive cross-entity interactions while strengthening the identity of each entity. To support systematic evaluation, we introduce SLIM (Semantic Leakage in IMages), the first dataset dedicated to semantic leakage, comprising 1,130 human-verified samples spanning diverse scenarios, together with a novel automatic evaluation framework. Experiments demonstrate that DeLeaker consistently outperforms all baselines, even when they are provided with external information, achieving effective leakage mitigation without compromising fidelity or quality. These results underscore the value of attention control and pave the way for more semantically precise T2I models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。