arXiv:2608.11343cs.AIcs.CV2026-08

对比大模型与原生多模态嵌入在图文检索中的表现

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

论文配图:Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
图 1 · 摘自论文原文
  • 用GPT-4.1和Claude Sonnet直接做图文检索,无需训练
  • 两者效果与Gemini Embedding 2相当,均在Flickr30k上表现优异
  • 预计算嵌入后,原生多模态模型更适合低延迟场景

跨文本、图像、视频和音频的多模态检索与分类传统上依赖双编码器模型,通过对比学习对齐视觉与文本表示。2026年3月发布的Gemini Embedding 2是谷歌首个原生多模态嵌入模型,可将文本、图像、视频、音频和文档映射到统一共享空间,引发多模态检索系统竞争。与此同时,前沿大语言模型(LLMs)也展现出强大的视觉理解能力,引发疑问:它们能否作为有效的零样本排序器?本研究首次在Flickr30k数据集上直接比较原生多模态嵌入与基于LLM的视觉排序。结果表明,GPT-4.1与Claude Sonnet 4.6的表现与Gemini Embedding 2相当。此外,一旦嵌入预先计算完成,原生多模态嵌入更适合低延迟应用场景。

原文摘要 · Abstract (English)

Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.

多模态图像检索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。