arXiv:2502.08468cs.CVcs.AI2025-02ACL被引 52

用高质量合成数据训练多模态多语言嵌入模型,性能领先

mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

  • 基于三标准合成跨模态对齐的多语言数据
  • 在MMEB和XTD基准上达到顶尖表现
  • 适合需要多语言多模态理解的研究者

多模态嵌入模型能将文本、图像等异构数据映射到统一表征空间,但标注数据稀缺制约其性能。现有方法虽利用数据合成缓解此问题,但合成数据质量仍是关键瓶颈。本文提出高质量合成数据的三大标准:覆盖广泛任务与模态、跨模态语义一致、保持真实细节。据此,我们构建了涵盖多种任务、模态组合与语言的合成数据集,通过单次多模态大模型深度推理生成,并结合真实图像与精准文本,通过自评估与迭代优化保证真实性。基于这些高质量数据,我们训练了多模态多语言E5模型mmE5。大量实验表明,mmE5在MMEB基准上达到当前最优,在XTD基准上展现卓越多语言能力。代码、数据集与模型已开源。

原文摘要 · Abstract (English)

Multimodal embedding models have gained significant attention for their ability to map data from different modalities, such as text and images, into a unified representation space. However, the limited labeled multimodal data often hinders embedding performance. Recent approaches have leveraged data synthesis to address this problem, yet the quality of synthetic data remains a critical bottleneck. In this work, we identify three criteria for high-quality synthetic multimodal data. First, broad scope ensures that the generated data covers diverse tasks and modalities, making it applicable to various downstream scenarios. Second, robust cross-modal alignment makes different modalities semantically consistent. Third, high fidelity ensures that the synthetic data maintains realistic details to enhance its reliability. Guided by these principles, we synthesize datasets that: (1) cover a wide range of tasks, modality combinations, and languages, (2) are generated via a deep thinking process within a single pass of a multimodal large language model, and (3) incorporate real-world images with accurate and relevant texts, ensuring fidelity through self-evaluation and refinement. Leveraging these high-quality synthetic and labeled datasets, we train a multimodal multilingual E5 model mmE5. Extensive experiments demonstrate that mmE5 achieves state-of-the-art performance on the MMEB Benchmark and superior multilingual performance on the XTD benchmark. Our codes, datasets and models are released in https://github.com/haon-chen/mmE5.

多模态多语言嵌入模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。