arXiv:2412.08802cs.CLcs.CV2024-12被引 42

jina-clip-v2提升多语言图文理解,支持纯文本与跨模态任务。

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

  • 多任务多阶段对比学习,融合文本对、三元组和图文对训练。
  • 在英/多语言零样本文本检索与语义相似度上超越现有CLIP模型。
  • 支持多语言(29种)及复杂文档图像理解,嵌入维度可调。

对比语言-图像预训练(CLIP)广泛用于跨模态信息检索和多模态理解任务,但其主要针对跨模态视觉-语言任务优化,在单模态文本任务中表现不佳,且通常仅在英文数据集上训练,缺乏多语言理解能力。此外,现有基于CLIP的模型对视觉丰富的文档理解不足。本文提出jina-clip-v2,一种通过多任务、多阶段对比学习范式,利用文本对、三元组和图像-文本对联合训练的对比视觉-语言模型,以支持纯文本与跨模态任务。模型采用多语言文本编码器,训练数据扩展至29种非英语语言(包括印地语、中文、德语、法语等)及视觉丰富的文档图像。评估表明,jina-clip-v2在英/多语言环境下,于零样本文本检索、语义文本相似度和跨模态检索任务中显著优于现有CLIP模型。该模型支持灵活的嵌入维度选择,用户可根据需求调整表示粒度。模型已开源,可通过Hugging Face获取。

原文摘要 · Abstract (English)

Contrastive Language-Image Pretraining (CLIP) has been widely used for crossmodal information retrieval and multimodal understanding tasks. However, CLIP models are mainly optimized for crossmodal vision-language tasks and underperform in single-mode text tasks. Moreover, these models are often trained on English datasets and therefore lack multilingual understanding. Additionally, from a visual understanding perspective, previous CLIP-based models exhibit insufficient understanding of visually rich documents. In this work, we propose jina-clip-v2, a contrastive vision-language model trained on text pairs, triplets and image-text pairs via a multi-task and multi-stage contrastive learning paradigm in order to support both text-only and crossmodal tasks. We employ a multilingual text encoder and expand the training dataset to include multilingual texts from 29 non-English languages, including Hindi, Chinese, German, French, and others, as well as images of visually rich documents. We evaluate the model's performance and show that jina-clip-v2 achieves notable improvements over state-of-the-art CLIP-based models in zero-shot text-only retrieval, semantic textual similarity, and crossmodal retrieval tasks in both English and multilingual settings. jina-clip-v2 also provides for flexibility in embedding dimensionality, enabling users to select the granularity of the representations. jina-clip-v2 is publicly available at https://huggingface.co/jinaai/jina-clip-v2.

多语言图文对齐嵌入模型CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。