arXiv:2605.27295cs.CV2026-05被引 5

Gemini Embedding 2实现多模态统一嵌入,支持图文音视频联合表征。

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

论文配图:Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
图 1 · 摘自论文原文
  • 基于Gemini多模态能力,统一处理跨模态混合输入。
  • 在MSCOCO等基准上达62.9 R@1,超越专用模型性能。
  • 零样本跨领域表现强,适合搜索、推荐等下游场景。

我们提出Gemini Embedding 2,一种原生多模态嵌入模型,可将视频、音频、图像和文本统一映射到同一表示空间。利用Gemini的多模态能力,该模型能对任意交错组合的多模态输入生成嵌入,并在多种任务中表现优异。通过大规模对比学习与多任务多阶段训练,其在关键嵌入基准上达到领先水平,涵盖单模态、跨模态及多模态检索任务。在MSCOCO上取得62.9 R@1,在Vatex上达68.8 NDCG@10,MTEB多语言与代码任务分别达69.9与84.0,均超越专用模型。其统一能力使其适用于RAG、推荐与搜索等下游应用。此外,其在天文学、生物科学、艺术与烹饪等领域的零样本表现稳健,具备强泛化性,是即插即用的可靠表征工具。

原文摘要 · Abstract (English)

We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage the multimodal capabilities of Gemini to produce embeddings for arbitrary combinations of interleaved inputs across all these modalities that generalize well across a wide variety of tasks. Applying large-scale contrastive learning in a multi-task multi-stage training setup, we achieve state-of-the-art performance on key embedding benchmarks including unimodal, cross-modal, and multimodal retrieval spanning a diverse set of tasks. We show that our embedding model demonstrates strong performance (with a score of 62.9 R@1 on MSCOCO, 68.8 NDCG@10 on Vatex, 69.9 on MTEB multilingual and 84.0 on MTEB Code) across a variety of tasks surpassing the performance of specialized models. These unified capabilities make Gemini Embedding 2 a promising candidate for downstream use cases such as RAG, recommendation and search. Furthermore, its robust zero-shot performance across distinct fields - from astronomy and bioscience to fine arts and the culinary arts - establishes it as a highly reliable, out-of-the-box representation even for specialized domains.

多模态嵌入统一表征零样本RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。