用分离的语义与感知描述文本,让图像特征更精细可控。
Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation
- 构建图文分离数据集,用感知与语义双文本引导特征解耦。
- 在CLIP基础上解耦出纯感知特征,提升图像质量评估精度。
- 适合需要精细控制图像感知质量的研究者使用。
对比视觉语言模型(如CLIP)在大规模I&1T(一张图配一句文)数据集上训练后,展现出优秀的零样本语义识别能力,但其多模态表示常混合语义与感知信息,偏重语义。这对图像质量评估(IQA)和条件图像生成(CIG)等任务不利,因这类任务需对感知与语义特征分别控制。为此,本文提出一种新的多模态解耦表示学习框架,利用分离的文本指导图像特征解耦。首先构建I&2T数据集(一张图配一句感知描述、一句语义描述),实现图文解耦。随后以这些分离文本为监督信号,从CLIP原始粗粒度特征空间中解耦出纯感知表示,命名为DeCLIP。最终,解耦特征用于图像质量评估(技术质量与审美质量)及条件图像生成。大量实验表明该方法在两项任务上均具优势。数据集、代码与模型将公开。
原文摘要 · Abstract (English)
Contrastive vision-language models, such as CLIP, have demonstrated excellent zero-shot capability across semantic recognition tasks, mainly attributed to the training on a large-scale I&1T (one Image with one Text) dataset. This kind of multimodal representations often blend semantic and perceptual elements, placing a particular emphasis on semantics. However, this could be problematic for popular tasks like image quality assessment (IQA) and conditional image generation (CIG), which typically need to have fine control on perceptual and semantic features. Motivated by the above facts, this paper presents a new multimodal disentangled representation learning framework, which leverages disentangled text to guide image disentanglement. To this end, we first build an I&2T (one Image with a perceptual Text and a semantic Text) dataset, which consists of disentangled perceptual and semantic text descriptions for an image. Then, the disentangled text descriptions are utilized as supervisory signals to disentangle pure perceptual representations from CLIP's original `coarse' feature space, dubbed DeCLIP. Finally, the decoupled feature representations are used for both image quality assessment (technical quality and aesthetic quality) and conditional image generation. Extensive experiments and comparisons have demonstrated the advantages of the proposed method on the two popular tasks. The dataset, code, and model will be available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。