用少量数据实现高质量多模态对齐,靠结构化正则化和深层相似性对齐。
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
- 通过保留单模态编码器隐空间邻域结构,设计新正则化方法。
- 仅需数万样本即达高精度,分类与检索任务提升超50%和90%。
- 适合数据稀缺领域,可嵌入现有方法快速增效。
多模态模型在需要多模态对齐的复杂任务(如零样本分类和跨模态检索)中表现强劲,但通常依赖数百万对齐样本,这在许多领域成本过高或不可行。本文探索在极有限配对数据下构建多模态模型的可行性,通过对齐预训练的单模态基础模型实现。结果显示,仅需数万对样本(不足领域常规用量的1%),即可实现高质量对齐。为此,我们提出STRUCTURE,一种有效保持单模态编码器隐空间邻域几何结构的正则化技术。此外,我们发现对齐最后一层往往效果不佳,而对齐跨模态表示相似度最高的层更具优势。这两项技术可无缝融入现有对齐方法,在24个零样本图像分类与检索基准上带来显著提升:分类任务平均相对改进51.6%,检索任务达91.8%。结果表明该框架在小样本多模态学习中高效且普适,为资源受限领域提供可行路径。
原文摘要 · Abstract (English)
Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, existing models typically rely on millions of paired multimodal samples, which are prohibitively expensive or infeasible to obtain in many domains. In this work, we explore the feasibility of building multimodal models with limited amount of paired data by aligning pretrained unimodal foundation models. We show that high-quality alignment is possible with as few as tens of thousands of paired samples$\unicode{x2013}$less than $1\%$ of the data typically used in the field. To achieve this, we introduce STRUCTURE, an effective regularization technique that preserves the neighborhood geometry of the latent space of unimodal encoders. Additionally, we show that aligning last layers is often suboptimal and demonstrate the benefits of aligning the layers with the highest representational similarity across modalities. These two components can be readily incorporated into existing alignment methods, yielding substantial gains across 24 zero-shot image classification and retrieval benchmarks, with average relative improvement of $51.6\%$ in classification and $91.8\%$ in retrieval tasks. Our results highlight the effectiveness and broad applicability of our framework for limited-sample multimodal learning and offer a promising path forward for resource-constrained domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。