arXiv:2504.10471cs.CVcs.CL2025-04ICCV被引 9

构建首个跨130项任务的图像嵌入评估基准,揭示模型在多任务中的真实表现。

MIEB: Massive Image Embedding Benchmark

  • 覆盖38种语言、130个任务,按8类高阶能力分组评估
  • 50个模型测试显示无单一模型通用于所有任务
  • 发现顶尖视觉模型能精准表征文本视觉内容,但对混杂信息匹配能力弱

图像表示通常通过孤立的、任务特定的评估协议进行,导致对模型能力的理解碎片化。例如,擅长图像聚类的模型是否也擅长基于文本检索相关图像尚不明确。我们提出大规模图像嵌入基准(MIEB),用于评估图像与图文嵌入模型在迄今最广泛范围内的性能。MIEB涵盖38种语言下的130项独立任务,分为8个高层次类别。我们在该基准上对50个模型进行了评测,发现没有任何一个方法在所有任务类别中占优。研究揭示了先进视觉模型的隐藏能力,如其对文本的准确视觉表征,以及在交错编码和存在混淆因子时匹配图文的局限性。此外,我们还发现视觉编码器在MIEB上的表现与它们在多模态大语言模型中的表现高度相关。代码、数据集和排行榜已公开于https://github.com/embeddings-benchmark/mteb。

原文摘要 · Abstract (English)

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.

图像嵌入多任务评估基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。