用集合思维提升少样本多模态学习效果
SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning

- 将多模态对齐视为集合优化,利用多个数据增强和描述捕捉跨模态结构
- 在仅需数万样本时达到强泛化性能,相比传统方法减少数量级数据需求
- 适合低数据场景下的多模态模型训练,尤其适用于标注成本高的任务
尽管多模态基础模型取得进展,但其依赖大规模成对数据限制了在低数据和罕见场景中的应用。关键瓶颈在于采用实例级建模,仅最大化图像-文本对间的相关性,忽略模态间潜在几何结构,导致模态差距。本文提出一种组合范式,引入子模态对齐器(SMA),将同一实体的多个增强和描述视为一个集合,利用多重正向关联捕捉更丰富的跨模态结构。SMA基于子模互信息(SMI)构建原则性目标,联合最大化跨模态互信息并降低交叉模态差异。该方法显著提升有限数据下的信息利用率。在CLIP基准的14个零样本分类与检索任务上评估,SMA在低数据条件下表现一致提升,仅用数万样本即实现强泛化,远少于标准方法所需。结果凸显集合建模与子模目标对高效多模态学习的重要性。
原文摘要 · Abstract (English)
Despite the recent success of Multimodal Foundation Models (FMs), their reliance on massive paired datasets limits their applicability in low-data and rare-scenario settings where aligned data is scarce and expensive. A key bottleneck is the adoption of an instance-level formulation, which learns alignment by maximizing correlation between individual image-text pairs while neglecting the underlying geometric structure across modalities resulting in a modality gap across input modalities. In this paper, we propose a combinatorial paradigm for multimodal alignment that moves beyond pairwise learning and introduce the \emph{Submodular Modality Aligner (SMA)}, which treats multiple augmentations and descriptions of an entity as a set, leveraging multiple descriptions of the data to capture richer cross-modal structure. We instantiate SMA using a principled objective based on Submodular Mutual Information (SMI), which jointly maximizes inter-modality mutual information while reducing cross-modal divergence. This formulation enables the model to effectively utilize multiple positive associations and extract significantly more information from limited data. We evaluate SMA on 14 zero-shot classification and retrieval tasks from the CLIP benchmark and demonstrate consistent gains in the low-data regime. Notably, SMA achieves strong multimodal generalization using only tens of thousands of samples. This is orders of magnitude fewer than standard approaches. Our results highlight the importance of set-based formulations and submodular objectives for data-efficient multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。