arXiv:2605.14270cs.CV2026-05中稿 · ICML被引 1

通过增强缺失概念信号,提升多模态扩散模型生成完整性。

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

论文配图:Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
图 1 · 摘自论文原文
  • 发现文本嵌入中存在可识别的缺失信号。
  • 在FLUX.1-Dev和SD3.5-Medium上显著减少概念遗漏。
  • 适合关注图像生成准确性的研究者使用。

多模态扩散变压器(MM-DiTs)在文生图任务中表现卓越,但常出现指定对象或属性未生成的问题。通过对文本词元进行线性探测,我们发现文本嵌入能区分表征目标概念缺失的‘缺失信号’。基于此,提出缺失信号干预(OSI)方法,通过放大该信号主动促进缺失概念的生成。在FLUX.1-Dev和SD3.5-Medium上的实验表明,OSI能显著缓解概念遗漏,即使在极端场景下也有效。

原文摘要 · Abstract (English)

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects or attributes fail to emerge in the generated image. By performing linear probing on text tokens, we demonstrate that text embeddings can distinguish a characteristic `omission signal' representing the absence of target concepts. Leveraging this insight, we propose Omission Signal Intervention (OSI), which amplifies the omission signal to actively catalyze the generation of missing concepts. Comprehensive experiments on FLUX.1-Dev and SD3.5-Medium demonstrate that OSI significantly alleviates concept omission even in extreme scenarios.

扩散模型文生图概念缺失多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。