arXiv:2602.06451cs.LG2026-02

突破数据集限制,实现跨数据集模态自由组合

BrokenBind: Universal Modality Exploration beyond Dataset Boundaries

  • 利用多数据集共享模态,生成伪嵌入填补缺失模态
  • 支持任意两模态绑定,无需数据集对齐
  • 适用于小样本场景,适合跨域多模态研究

多模态学习通过融合多种模态来全面理解现实问题。现有方法通常在特定联合嵌入空间中直接绑定模态,但其能力受限于给定数据集中的模态,导致在下游任务中面对未出现的模态时存在偏差。由于这种僵化性,传统方法的实用性受制于多模态数据获取成本。本文提出BrokenBind,专注于从不同数据集中绑定目标模态。该方法同时利用包含目标模态的多个数据集及一个共享模态,尽管数据集间因分布差异无法直接对应,仍可通过捕捉其关系生成伪嵌入,填补缺失模态,实现灵活且泛化的多模态学习。在该框架下,任意两模态可自由绑定,摆脱数据集限制,实现通用模态探索。进一步研究了需超过两个数据集进行模态绑定的增强场景,并验证了BrokenBind在低数据条件下的有效性。大量实验表明,其性能优于知名多模态基线方法。

原文摘要 · Abstract (English)

Multi-modal learning combines various modalities to provide a comprehensive understanding of real-world problems. A common strategy is to directly bind different modalities together in a specific joint embedding space. However, the capability of existing methods is restricted within the modalities presented in the given dataset, thus they are biased when generalizing to unpresented modalities in downstream tasks. As a result, due to such inflexibility, the viability of previous methods is seriously hindered by the cost of acquiring multi-modal datasets. In this paper, we introduce BrokenBind, which focuses on binding modalities that are presented from different datasets. To achieve this, BrokenBind simultaneously leverages multiple datasets containing the modalities of interest and one shared modality. Though the two datasets do not correspond to each other due to distribution mismatch, we can capture their relationship to generate pseudo embeddings to fill in the missing modalities of interest, enabling flexible and generalized multi-modal learning. Under our framework, any two modalities can be bound together, free from the dataset limitation, to achieve universal modality exploration. Further, to reveal the capability of our method, we study intensified scenarios where more than two datasets are needed for modality binding and show the effectiveness of BrokenBind in low-data regimes. Through extensive evaluation, we carefully justify the superiority of BrokenBind compared to well-known multi-modal baseline methods.

多模态学习跨数据集模态绑定低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。