新基准ARK评测多模态检索在专业知识与复杂推理上的表现
ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
- 从知识领域与推理能力双维度构建评测框架
- 16类视觉数据+硬负样本,检验多步推理能力
- 发现细粒度视觉推理是当前模型主要瓶颈
现有多模态检索基准多聚焦日常图像的语义匹配,缺乏对专业领域知识和复杂推理能力的诊断。为此,我们提出ARK,一个从两个互补视角评估多模态检索的基准:(i) 知识领域(5个领域,17种子类型),刻画检索依赖的内容与专业性;(ii) 推理技能(6类),刻画识别正确候选所需对多模态证据的推理类型。ARK涵盖16种异构视觉数据类型,支持单模态与多模态查询与候选,并采用针对性强的硬负样本以避免捷径匹配。我们在ARK上评估了23个代表性文本与多模态检索器,发现知识密集型与推理密集型检索之间存在显著差距,细粒度视觉与空间推理成为持续性瓶颈。此外,简单的重排序与改写可带来一致提升,但仍有巨大改进空间。
原文摘要 · Abstract (English)
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 23 representative text-based and multimodal retrievers on ARK and observe a pronounced gap between knowledge-intensive and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning emerging as persistent bottlenecks. We further show that simple enhancements such as re-ranking and rewriting yield consistent improvements, but substantial headroom remains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。