提出新方法检测主体驱动生成中的视觉不一致,可定位问题区域。
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- 通过对比架构分离扩散模型的语义与视觉特征
- 构建新指标VSM,精准量化并定位生成不一致处
- 首次实现对生成不一致的量化与空间定位,适合图像质量评估
我们提出一种新方法,从预训练扩散模型的主干网络中解耦视觉与语义特征,实现类似语义对应关系的视觉对应。尽管扩散模型主干已知蕴含丰富语义信息,但为支持图像生成,也必须包含视觉特征,然而因缺乏标注数据,分离这些视觉特征极具挑战。为此,我们设计了一套自动化流程,基于现有主体驱动图像生成数据集构建带有语义与视觉对应标注的图像对,并提出对比架构以分离两类特征。利用解耦表示,我们提出了新的视觉语义匹配(VSM)指标,用于量化主体驱动图像生成中的视觉不一致。实验证明,该方法在量化视觉不一致方面优于基于全局特征的指标如CLIP、DINO及视觉-语言模型,同时具备不一致区域的空间定位能力。据我们所知,这是首个支持不一致量化与定位的方法,为推进该任务提供了重要工具。
原文摘要 · Abstract (English)
We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While diffusion model backbones are known to encode semantically rich features, they must also contain visual features to support their image synthesis capabilities. However, isolating these visual features is challenging due to the absence of annotated datasets. To address this, we introduce an automated pipeline that constructs image pairs with annotated semantic and visual correspondences based on existing subject-driven image generation datasets, and design a contrastive architecture to separate the two feature types. Leveraging the disentangled representations, we propose a new metric, Visual Semantic Matching (VSM), that quantifies visual inconsistencies in subject-driven image generation. Empirical results show that our approach outperforms global feature-based metrics such as CLIP, DINO, and vision--language models in quantifying visual inconsistencies while also enabling spatial localization of inconsistent regions. To our knowledge, this is the first method that supports both quantification and localization of inconsistencies in subject-driven generation, offering a valuable tool for advancing this task. Project Page:https://abdo-eldesokey.github.io/mind-the-glitch/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。