arXiv:2508.20322cs.CV2025-08被引 1

通过稀疏线性概念子空间分离视觉语义嵌入,实现更精准的概念过滤检索。

Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)

  • 用分组字典学习构建稀疏非负组合的语义组件向量
  • 在多标签数据下提升概念过滤图像检索精度
  • 适用于压缩嵌入与自监督模型,支持零样本标注

视觉语言共嵌入网络(如CLIP)提供了包含语义信息的潜在嵌入空间,可用于下游任务。我们假设该嵌入空间可通过分解为多个概念特定的分量向量来解耦,这些向量位于不同子空间中。本文提出一种监督字典学习方法,估计一个由字典(原子)的稀疏非负组合构成的线性合成模型,其分组激活匹配多标签信息。每个概念分量是非负组合的原子集合,对应某一标签。通过一种保证收敛的交替优化策略优化具有结构的字典。利用文本共嵌入,我们基于单词嵌入中最佳逼近某概念原子组的词,找到语义有意义的描述;无监督字典学习可借助概念标签的文本嵌入,对训练集图像进行零样本分类,生成实例级多标签。实验表明,由稀疏线性概念子空间(SLiCS)提供的解耦嵌入可实现更精确的概念过滤图像检索(及基于图像到提示的条件生成)。我们还将SLiCS应用于TiTok的高压缩自编码器嵌入和自监督DINOv2的嵌入。定量与定性结果均显示,所有嵌入下的概念过滤检索精度均有提升。

原文摘要 · Abstract (English)

Vision-language co-embedding networks, such as CLIP, provide a latent embedding space with semantic information that is useful for downstream tasks. We hypothesize that the embedding space can be disentangled to separate the information on the content of complex scenes by decomposing the embedding into multiple concept-specific component vectors that lie in different subspaces. We propose a supervised dictionary learning approach to estimate a linear synthesis model consisting of sparse, non-negative combinations of groups of vectors in the dictionary (atoms), whose group-wise activity matches the multi-label information. Each concept-specific component is a non-negative combination of atoms associated to a label. The group-structured dictionary is optimized through a novel alternating optimization with guaranteed convergence. Exploiting the text co-embeddings, we detail how semantically meaningful descriptions can be found based on text embeddings of words best approximated by a concept's group of atoms, and unsupervised dictionary learning can exploit zero-shot classification of training set images using the text embeddings of concept labels to provide instance-wise multi-labels. We show that the disentangled embeddings provided by our sparse linear concept subspaces (SLiCS) enable concept-filtered image retrieval (and conditional generation using image-to-prompt) that is more precise. We also apply SLiCS to highly-compressed autoencoder embeddings from TiTok and the latent embedding from self-supervised DINOv2. Quantitative and qualitative results highlight the improved precision of the concept-filtered image retrieval for all embeddings.

嵌入解耦概念分离图像检索字典学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。