新基准挑战视觉检索极限,推动模型在复杂场景下的性能突破
ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
- 通过盲式上下文查询、长文本跨文档查询提升测试难度
- 多语言数据集上初始模型表现仍有显著提升空间
- 支持社区共建,持续更新以保持评测前沿性
ViDoRe Benchmark V1 已接近性能饱和,顶尖模型 nDCG@5 超过 90%,难以有效区分改进。ViDoRe Benchmark V2 通过盲式上下文查询、长且跨文档的查询以及混合合成与人工参与的查询生成方式,引入更真实、更具挑战性的检索场景。该基准包含四个多样化、多语言的数据集,并提供清晰的评估说明。初步结果显示模型仍有巨大提升空间,并揭示了模型泛化能力和多语言表现的洞见。本基准设计为可持续演进的资源,鼓励社区贡献,以确保未来评估的相关性。
原文摘要 · Abstract (English)
The ViDoRe Benchmark V1 was approaching saturation with top models exceeding 90% nDCG@5, limiting its ability to discern improvements. ViDoRe Benchmark V2 introduces realistic, challenging retrieval scenarios via blind contextual querying, long and cross-document queries, and a hybrid synthetic and human-in-the-loop query generation process. It comprises four diverse, multilingual datasets and provides clear evaluation instructions. Initial results demonstrate substantial room for advancement and highlight insights on model generalization and multilingual capability. This benchmark is designed as a living resource, inviting community contributions to maintain relevance through future evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。