让AI生成图像时主动忽略特定物体,效果更准且不破坏画面质量。
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

- 在注意力特征空间中对负向提示做正交处理,只移除干扰内容
- 在FLUX模型上实现90%以上概念抑制率,图像质量无明显下降
- 支持多个概念同时屏蔽,可调节抑制强度,适合精准控图场景
文本到图像(T2I)模型生成能力日益强大,但明确排除特定物体或属性仍是根本性挑战。现有方法如提示词否定、后期编辑和负向引导,在显式概念抑制方面仍不足,常无法完全移除目标概念或降低图像质量。为此,我们提出一种训练无关的方法——注意力特征空间中的正交负向引导,该方法作用于基于MM-DiT的T2I变换器的注意力输出空间。通过将负向提示注意力特征与正向提示特征正交化,并仅减去正交分量,实现对不需要概念的有效抑制,同时保留期望语义。在FLUX-dev和FLUX-schnell上的实验表明,该方法在概念抑制、提示对齐与图像质量之间取得良好权衡。人工评估中,相比第二优基线提升18.78%。进一步验证了方法支持多概念抑制及可调节抑制强度的能力。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fundamentally challenging problem. Existing approaches, including prompt negation, post-hoc editing, and negative guidance, remain insufficient for explicit concept suppression, often failing to remove the target concept or degrading overall image quality. To this end, we propose Orthogonal Negative Guidance in attention feature space, a training-free method that operates in the attention output space of MM-DiT-based T2I transformers. Our method orthogonalizes negative-prompt attention features with respect to positive-prompt features and subtracts only the orthogonal component, suppressing unwanted concepts while preserving desired semantics. Experiments on FLUX-dev and FLUX-schnell show that our method achieves favorable trade-offs between concept suppression, prompt alignment, and image quality. In human evaluation, our method outperforms the second-best baseline by 18.78%. We further show that our method supports multi-concept suppression and adjustable concept suppression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。