arXiv:2505.05635cs.CV2025-05被引 1

用百科知识增强图像识别,让模型在成千上万鸟种中更准地认出物种。

Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation

  • 结合百科摘要与视觉信息,动态重排序候选物种。
  • 在五项鸟类基准上使顶尖模型准确率提升18.0%。
  • 适合需要持续更新物种库的生态监测系统使用。

开放词汇物种识别是计算机视觉中的重大挑战,尤其在鸟类学领域,新类群不断被发现。尽管CUB-200-2011和Birdsnap等基准推动了封闭词汇下的细粒度识别,但难以应对真实场景。我们发现当前系统在含数千候选物种的开放词汇设置下性能下降超过30%,主要因视觉相似和语义模糊的干扰项增多。为此,提出视觉重排序检索增强生成(VR-RAG)框架,将结构化百科知识与识别任务结合。从11,202种鸟类的维基百科文章中提炼出简洁、具有区分性的摘要,并基于这些摘要检索候选物种。与以往仅依赖文本的方法不同,VR-RAG在检索阶段融合视觉信息,确保最终预测既与文本描述相关,又与查询图像视觉一致。在五个鸟类分类基准及两个额外领域上的大量实验表明,该方法使最先进的Qwen2.5-VL模型平均性能提升18.0%。

原文摘要 · Abstract (English)

Open-vocabulary species recognition is a major challenge in computer vision, particularly in ornithology, where new taxa are continually discovered. While benchmarks like CUB-200-2011 and Birdsnap have advanced fine-grained recognition under closed vocabularies, they fall short of real-world conditions. We show that current systems suffer a performance drop of over 30\% in realistic open-vocabulary settings with thousands of candidate species, largely due to an increased number of visually similar and semantically ambiguous distractors. To address this, we propose Visual Re-ranking Retrieval-Augmented Generation (VR-RAG), a novel framework that links structured encyclopedic knowledge with recognition. We distill Wikipedia articles for 11,202 bird species into concise, discriminative summaries and retrieve candidates from these summaries. Unlike prior text-only approaches, VR-RAG incorporates visual information during retrieval, ensuring final predictions are both textually relevant and visually consistent with the query image. Extensive experiments across five bird classification benchmarks and two additional domains show that VR-RAG improves the average performance of the state-of-the-art Qwen2.5-VL model by 18.0%.

物种识别视觉语言模型知识增强开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。