通过智能筛选模态数据与自适应训练,提升多模态检索效果
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- 基于模态特性筛选数据并设计针对性训练策略
- 在多个基准上达到当前最佳性能,超越现有方法
- 适合需要构建鲁棒多模态系统的研究人员
多模态信息检索(MIR)面临数据源异构性与跨模态对齐复杂性的挑战。尽管先前研究已发现特征空间中的模态差异,但系统性解决方案仍待探索。本文提出UNITE框架,通过数据筛选与模态感知训练配置,首次全面分析模态特异性数据属性对下游任务性能的影响。我们引入模态感知掩码对比学习(MAMCL),缓解不同模态实例间的竞争关系。在多个多模态检索基准上,该框架显著优于现有方法。大量实验证明,有策略的模态筛选与定制化训练协议对跨模态表示学习至关重要。本工作不仅提升了MIR性能,也为未来多模态系统研究提供基础范式。项目主页:https://friedrichor.github.io/projects/UNITE。
原文摘要 · Abstract (English)
Multimodal information retrieval (MIR) faces inherent challenges due to the heterogeneity of data sources and the complexity of cross-modal alignment. While previous studies have identified modal gaps in feature spaces, a systematic approach to address these challenges remains unexplored. In this work, we introduce UNITE, a universal framework that tackles these challenges through two critical yet underexplored aspects: data curation and modality-aware training configurations. Our work provides the first comprehensive analysis of how modality-specific data properties influence downstream task performance across diverse scenarios. Moreover, we propose Modal-Aware Masked Contrastive Learning (MAMCL) to mitigate the competitive relationships among the instances of different modalities. Our framework achieves state-of-the-art results on multiple multimodal retrieval benchmarks, outperforming existing methods by notable margins. Through extensive experiments, we demonstrate that strategic modality curation and tailored training protocols are pivotal for robust cross-modal representation learning. This work not only advances MIR performance but also provides a foundational blueprint for future research in multimodal systems. Our project is available at https://friedrichor.github.io/projects/UNITE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。