arXiv:2506.18902cs.AIcs.CL2025-06被引 98

38亿参数多模态模型,支持图文统一检索与复杂视觉内容理解。

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval

  • 采用晚期交互架构,支持单向量与多向量嵌入统一表示。
  • 在跨模态检索中达到顶尖水平,尤其擅长表格、图表等视觉内容。
  • 附带新基准Jina-VDR,专测复杂视觉信息检索能力。

我们提出 jina-embeddings-v4,一个拥有38亿参数的多模态嵌入模型,通过创新架构实现文本与图像表征的统一,支持单向量和多向量嵌入的晚交互方式。模型引入任务特定的低秩适配(LoRA)适配器,以优化在查询-文档检索、语义文本相似性及代码搜索等多种检索场景下的表现。全面评估表明,jina-embeddings-v4 在单模态与跨模态检索任务上均达到当前最优性能,尤其在处理包含表格、图表、示意图及混合媒体格式的视觉丰富内容时表现突出。为评估该能力,我们还推出了 Jina-VDR,一个专门针对视觉丰富图像检索设计的新基准。

原文摘要 · Abstract (English)

We introduce jina-embeddings-v4, a 3.8 billion parameter multimodal embedding model that unifies text and image representations through a novel architecture supporting both single-vector and multi-vector embeddings in the late interaction style. The model incorporates task-specific Low-Rank Adaptation (LoRA) adapters to optimize performance across diverse retrieval scenarios, including query-document retrieval, semantic text similarity, and code search. Comprehensive evaluations demonstrate that jina-embeddings-v4 achieves state-of-the-art performance on both single-modal and cross-modal retrieval tasks, with particular strength in processing visually rich content such as tables, charts, diagrams, and mixed-media formats. To facilitate evaluation of this capability, we also introduce Jina-VDR, a novel benchmark specifically designed for visually rich image retrieval.

多模态嵌入模型检索视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。