首个需要深度推理的跨学科多模态检索基准,逼真模拟专家级任务。
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- 构建23个领域1502个需深度推理的图文混合查询,由专家验证正例。
- 引入矛盾检索新任务,要求识别图像中冲突概念,提升推理难度。
- 支持多图查询与图文交错文档,更贴近真实复杂检索场景。
我们提出MRMR,首个面向专家级别的多学科多模态检索基准,需深度推理。该基准包含1,502个跨23个领域的查询,正样本由人类专家严格验证。相比以往基准,MRMR实现三大突破:其一,覆盖多元专业领域,支持细粒度模型域间对比;其二,查询具强推理性,如需解读显微镜切片图像,并引入“矛盾检索”新任务,要求模型识别概念冲突;其三,查询与文档采用图像-文本交错序列结构,不同于单一图像或单模态文档,更贴近真实多模态检索场景。我们在MRMR上评估4类多模态检索系统及14个前沿模型。结果显示,结合大模型生成图像描述的Qwen3-Embedding在文本嵌入中表现最佳,表明多模态检索仍有巨大提升空间。尽管最新模型如Ops-MM-Embedding在专家领域查询上表现尚可,但在推理密集型任务中仍显不足。我们认为MRMR将推动多模态检索向更真实、更具挑战性的方向发展。
原文摘要 · Abstract (English)
We introduce MRMR, the first expert-level multidisciplinary multimodal retrieval benchmark requiring intensive reasoning. MRMR contains 1,502 queries spanning 23 domains, with positive documents carefully verified by human experts. Compared to prior benchmarks, MRMR introduces three key advancements. First, it challenges retrieval systems across diverse areas of expertise, enabling fine-grained model comparison across domains. Second, queries are reasoning-intensive, with images requiring deeper interpretation such as diagnosing microscopic slides. We further introduce Contradiction Retrieval, a novel task requiring models to identify conflicting concepts. Finally, queries and documents are constructed as image-text interleaved sequences. Unlike earlier benchmarks restricted to single images or unimodal documents, MRMR offers a realistic setting with multi-image queries and mixed-modality corpus documents. We conduct an extensive evaluation of 4 categories of multimodal retrieval systems and 14 frontier models on MRMR. The text embedding model Qwen3-Embedding with LLM-generated image captions achieves the highest performance, highlighting substantial room for improving multimodal retrieval models. Although latest multimodal models such as Ops-MM-Embedding perform competitively on expert-domain queries, they fall short on reasoning-intensive tasks. We believe that MRMR paves the way for advancing multimodal retrieval in more realistic and challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。