通过优化数据对与模型结构,显著缩小跨模态语义差距。
Multimodal Data Curation Through Ranked Retrieval

- 用SNS筛选训练数据中互证最强的片段,减少噪声。
- 用EEE融合多专家嵌入,使跨模态相似度提升超90%。
- 适合大规模多模态数据清洗与高质量数据集构建者。
共享嵌入空间广泛用于多模态搜索与数据整理。但实践中常面临两个问题:一是嵌入偏向模态而非语义,导致内容匹配时仍按输入类型聚类;二是训练所用配对监督常含噪声。当混合多种异构、人工标注数据集时,这些问题相互加剧,损害跨模态检索效果。本文提出框架,同时优化训练对与嵌入模型。对称核采样(SNS)通过修剪原始输入与标注中互证最弱的部分,精炼训练对。专家嵌入引擎(EEE)利用学习的投影网络融合互补嵌入专家,并采用偏置感知目标,降低嵌入空间中由模态驱动的分离。实验表明,该方法平均将模态差距缩小超过90%,且作为数据整理工具表现优异,其生成的数据混编在下游模型性能上优于分层采样与传统整理基线。
原文摘要 · Abstract (English)
Shared embedding spaces are widely used for multimodal search and data curation. In practice, two problems often limit how well this works. First, embeddings can reflect modality more than meaning, so examples cluster by input type even when the underlying content matches. Second, the paired supervision used to train these spaces is often noisy. When we blend many heterogeneous, human-labeled datasets, these issues reinforce each other and degrade cross-modal retrieval. We present a framework that improves alignment by acting on both the training pairs and the embedding model. Symmetric Nucleus Subsampling (SNS) refines training pairs by trimming raw inputs and annotations to the portions that best support each other. Expert Embedding Engine (EEE) combines complementary embedding experts using a learned projection network, together with a bias-aware objective that reduces modality-driven separation in the embedding space. We demonstrate that this approach collapses the modality gap by over 90% on average vs base embedding experts and is a strong data curator, with datablends from our method outperforming stratified sampling and traditional curation baselines in downstream model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。