用相对表示学习细粒度跨模态对齐,小样本下效果更优
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data

- 通过可学习锚点捕捉跨模态的令牌级相似关系
- 在零样本分类、检索和分割上显著超越现有方法
- 适合数据稀缺场景下的跨模态模型对齐
多模态预训练表现出强泛化能力,但在配对数据稀缺领域不切实际。一种替代方案是事后多模态对齐,即用少量配对样本对已独立预训练的单模态编码器进行对齐。然而,现有方法主要关注全局表示对齐,忽略了片段-令牌间的结构关系,可能阻碍需细粒度跨模态匹配的任务迁移。为此,我们提出一种事后对齐方法,通过相对表示学习令牌级跨模态结构。具体地,将图像和文本表示为与各模态空间中可学习锚点的令牌级相似性,这些锚点被训练以对匹配对诱导一致的跨模态相似模式。尽管仅学习锚点而无需复杂投影层,该方法在零样本分类、跨模态检索和零样本分割任务上均显著优于现有方法。这凸显了在有限配对数据下建模细粒度跨模态结构的重要性。
原文摘要 · Abstract (English)
Multimodal pre-training demonstrates strong generalization performance, but this paradigm is often impractical in domains where paired data are scarce. A promising alternative is post-hoc multimodal alignment, which aligns separately pre-trained unimodal encoders using a limited number of paired examples. However, existing methods focus primarily on aligning global representations, missing patch-token relations. This may hinder transfer to tasks that require fine-grained cross-modal matching beyond coarse sample-level semantics. To address this issue, we propose a post-hoc alignment method that learns token-level cross-modal structure using relative representations. Specifically, we represent images and texts through their token-level similarities to a set of learnable anchors in each modality space, which are trained to induce consistent cross-modal similarity patterns for matched pairs. Despite learning only the anchors without heavy projection layers, our approach consistently outperforms existing methods in zero-shot classification, cross-modal retrieval, and zero-shot segmentation by a substantial margin. This highlights the importance of modeling fine-grained cross-modal structure for effective post-hoc multimodal alignment with limited paired data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。