用傅里叶分析分离图像结构与风格,提升视觉语言模型少样本泛化能力。
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
- 通过相位与幅度谱分别提取结构与风格特征,实现视觉表征解耦。
- 在15个数据集上验证,显著提升少样本学习性能。
- 适合研究视觉语言对齐与模型泛化性的研究人员。
大规模预训练视觉语言模型(VLMs)展现出强大的少样本学习能力。然而,这些方法通常学习整体性表征,导致图像的领域不变结构与领域特定风格隐式耦合。本文提出傅里叶感知表征学习(FARL),通过傅里叶分析显式解耦视觉表征。核心是双交叉注意力机制,可学习的表征标记分别查询图像的结构特征(来自相位谱)和风格特征(来自幅度谱)。该过程生成更丰富的解耦表征,并通过非对称注入策略深度嵌入VLM编码器,引导模型适应。大量实验在15个数据集上验证了方法的有效性。
原文摘要 · Abstract (English)
Large-scale pre-trained Vision-Language Models (VLMs) have demonstrated strong few-shot learning capabilities. However, these methods typically learn holistic representations where an image's domain-invariant structure is implicitly entangled with its domain-specific style. This presents an opportunity to further enhance generalization by disentangling these visual cues. In this paper, we propose Fourier-Attentive Representation Learning (FARL), a novel framework that addresses this by explicitly disentangling visual representations using Fourier analysis. The core of our method is a dual cross-attention mechanism, where learnable representation tokens separately query an image's structural features (from the phase spectrum) and stylistic features (from the amplitude spectrum). This process yields enriched, disentangled tokens that are then injected deep into the VLM encoders to guide adaptation. Our design, which includes an asymmetric injection strategy, forces the model to learn a more robust vision-language alignment. Extensive experiments on 15 datasets demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。