arXiv:2410.08469cs.LGcs.CL2024-10EMNLP被引 5

让CLIP的文本编码更可解释、可控,按语义重要性重加权单词。

Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP

  • 根据上下文重要性动态调整词元权重,改进文本编码
  • 在少样本图像分类与偏好检索中提升性能
  • 适合需要解释性和用户可控性的视觉语言应用

像CLIP这样的视觉-语言模型中的文本编码器,将自然语言输入映射到与图像共享的嵌入空间,使视觉任务能通过自然语言进行可解释分析。尽管句子中不同文本元素的重要性随上下文变化,但现有方法缺乏对这种重要性差异的建模。我们提出语义词元重加权框架SToRI,构建可解释且可控的文本嵌入。SToRI通过基于上下文重要性差异化地重加权语义成分,优化CLIP的文本编码过程,实现对强调内容的数据驱动和用户偏好响应式控制。在面向用户偏好的少样本图像分类与图像检索任务中,通过全面实验验证了SToRI的有效性。

原文摘要 · Abstract (English)

A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual elements within a sentence depending on the context, efforts to account for variation of importance in constructing text embeddings have been lacking. We propose a framework of Semantic Token Reweighting to build Interpretable text embeddings (SToRI), which incorporates controllability as well. SToRI refines the text encoding process in CLIP by differentially weighting semantic elements based on contextual importance, enabling finer control over emphasis responsive to data-driven insights and user preferences. The efficacy of SToRI is demonstrated through comprehensive experiments on few-shot image classification and image retrieval tailored to user preferences.

CLIP可解释性文本编码可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。