arXiv:2603.08942cs.CVcs.AI2026-03中稿 · Domain Generalizat…被引 1

用少量样本对齐跨域图像特征,提升视觉语言模型的零样本适应能力。

BiCLIP: Domain Canonicalization via Structured Geometric Transformation

  • 通过几何变换建模不同领域间的特征关系
  • 在11个基准上实现最优少样本分类性能
  • 方法简单高效,适合快速适配新领域

近年来,视觉语言模型(VLMs)展现出强大的零样本能力,但将其适配到特定领域仍面临挑战。基于最新理论发现——独立训练的VLMs之间存在可归一化的变换关系,我们进一步将这一思想扩展至领域层面。假设不同领域的图像特征由一种可归一化的几何变换关联,可通过少量锚点样本恢复该变换。少样本分类天然适合此设定,因有限标注样本可作为估计变换所需的锚点。受此启发,我们提出BiCLIP框架,对多模态特征施加定向变换以增强跨模态对齐。该方法具有极简结构和低参数量。在包括EuroSAT、DTD、FGVCAircraft在内的11个标准基准上的广泛评估表明,BiCLIP始终达到领先性能。此外,我们通过分析学习到变换的正交性和角度分布,实证验证了现有几何结论,确认结构化对齐是鲁棒领域适应的关键。代码已公开于https://github.com/QuantitativeImagingLaboratory/BilinearCLIP。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insights suggesting that independently trained VLMs are related by a canonical transformation, we extend this understanding to the concept of domains. We hypothesize that image features across disparate domains are related by a canonicalized geometric transformation that can be recovered using a small set of anchors. Few-shot classification provides a natural setting for this alignment, as the limited labeled samples serve as the anchors required to estimate this transformation. Motivated by this hypothesis, we introduce BiCLIP, a framework that applies a targeted transformation to multimodal features to enhance cross-modal alignment. Our approach is characterized by its extreme simplicity and low parameter footprint. Extensive evaluations across 11 standard benchmarks, including EuroSAT, DTD, and FGVCAircraft, demonstrate that BiCLIP consistently achieves state-of-the-art results. Furthermore, we provide empirical verification of existing geometric findings by analyzing the orthogonality and angular distribution of the learned transformations, confirming that structured alignment is the key to robust domain adaptation. Code is available at https://github.com/QuantitativeImagingLaboratory/BilinearCLIP

视觉语言模型少样本学习领域自适应几何变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。