构建自然世界图文检索基准,挑战AI理解生态细节能力
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
- 设计包含500万图像的iNat24数据集,匹配3.3万精确标签
- 专家级查询下最佳模型mAP@50仍不足50%,体现巨大挑战
- 面向生态研究需求,适合推动生物多样性智能检索系统研发
我们提出INQUIRE,一个面向多模态视觉语言模型的自然世界文本到图像检索基准,用于测试其在专家级查询下的表现。INQUIRE引入iNaturalist 2024(iNat24),一个包含五百万自然世界图像的新数据集,搭配250个专家级检索查询,所有相关图像均被全面标注,共33,000个匹配项。查询涵盖物种识别、环境、行为与外观等类别,强调需要细致图像理解与领域知识的任务。该基准评估两项核心任务:(1) INQUIRE-Fullrank,全数据集排名任务;(2) INQUIRE-Rerank,对前100项结果进行重排序。对多种近期多模态模型的评估显示,当前最佳模型在mAP@50上仍无法超过50%。此外,使用更强的多模态模型进行重排序可提升性能,但仍存在显著提升空间。通过聚焦科学驱动的生态挑战,INQUIRE旨在弥合人工智能能力与真实科研需求之间的差距,推动能助力生态与生物多样性研究的检索系统发展。数据集与代码已公开于https://inquire-benchmark.github.io
原文摘要 · Abstract (English)
We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along with 250 expert-level retrieval queries. These queries are paired with all relevant images comprehensively labeled within iNat24, comprising 33,000 total matches. Queries span categories such as species identification, context, behavior, and appearance, emphasizing tasks that require nuanced image understanding and domain expertise. Our benchmark evaluates two core retrieval tasks: (1) INQUIRE-Fullrank, a full dataset ranking task, and (2) INQUIRE-Rerank, a reranking task for refining top-100 retrievals. Detailed evaluation of a range of recent multimodal models demonstrates that INQUIRE poses a significant challenge, with the best models failing to achieve an mAP@50 above 50%. In addition, we show that reranking with more powerful multimodal models can enhance retrieval performance, yet there remains a significant margin for improvement. By focusing on scientifically-motivated ecological challenges, INQUIRE aims to bridge the gap between AI capabilities and the needs of real-world scientific inquiry, encouraging the development of retrieval systems that can assist with accelerating ecological and biodiversity research. Our dataset and code are available at https://inquire-benchmark.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。