扩散分类器能理解视觉组合性,但效果受数据域和时间步权重影响。
Diffusion Classifiers Understand Compositionality, but Conditions Apply

- 用扩散模型生成图像并转为分类器,测试其组合理解能力。
- 在10个数据集30多个任务中验证,发现性能与数据域密切相关。
- 首次揭示SD3-m对领域差异敏感,时间步加权可提升效果。
理解视觉场景是人类智能的核心。尽管判别模型推动了计算机视觉发展,却常难以处理组合性问题。而近期的文本到图像扩散模型在合成复杂场景方面表现出色,暗示其具备内在的组合能力。基于此,零样本扩散分类器被提出以复用扩散模型完成判别任务。尽管先前工作在组合判别任务中展现潜力,但受限于少量基准和浅层分析。为此,本文对扩散分类器在广泛组合任务中的判别能力进行了全面研究:涵盖三个扩散模型(SD 1.5、2.0 和首次引入的 3-m),覆盖10个数据集及超过30项任务。我们进一步探究目标数据集领域对性能的影响,为隔离领域效应,提出新诊断基准 extsc{Self-Bench},由扩散模型自身生成图像构成。最后,我们分析时间步加权的重要性,发现领域差距与时间步敏感性之间存在关联,尤其在SD3-m上显著。综上,扩散分类器理解组合性,但条件决定成败。代码与数据集见 https://github.com/eugene6923/Diffusion-Classifiers-Compositionality。
原文摘要 · Abstract (English)
Understanding visual scenes is fundamental to human intelligence. While discriminative models have significantly advanced computer vision, they often struggle with compositional understanding. In contrast, recent generative text-to-image diffusion models excel at synthesizing complex scenes, suggesting inherent compositional capabilities. Building on this, zero-shot diffusion classifiers have been proposed to repurpose diffusion models for discriminative tasks. While prior work offered promising results in discriminative compositional scenarios, these results remain preliminary due to a small number of benchmarks and a relatively shallow analysis of conditions under which the models succeed. To address this, we present a comprehensive study of the discriminative capabilities of diffusion classifiers on a wide range of compositional tasks. Specifically, our study covers three diffusion models (SD 1.5, 2.0, and, for the first time, 3-m) spanning 10 datasets and over 30 tasks. Further, we shed light on the role that target dataset domains play in respective performance; to isolate the domain effects, we introduce a new diagnostic benchmark \textsc{Self-Bench} comprised of images created by diffusion models themselves. Finally, we explore the importance of timestep weighting and uncover a relationship between domain gap and timestep sensitivity, particularly for SD3-m. To sum up, diffusion classifiers understand compositionality, but conditions apply! Code and dataset are available at https://github.com/eugene6923/Diffusion-Classifiers-Compositionality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。