用图文描述+外部知识增强,让AI更好识别罕见节肢动物。
Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification
- 结合图像描述与外部文本检索,提升分类可解释性。
- 在家族和属级分类上准确率显著优于普通大模型。
- 适合需要精准识别稀有物种的生物多样性研究者。
面对气候变化与节肢动物多样性急剧下降的挑战,基于生物图像的自动化分类成为研究热点。传统基于CNN或ViT的深度学习方法在长尾类别上表现不佳,且缺乏推理能力。本文将图像描述生成与检索增强生成(RAG)结合大语言模型(LLM),提升生物多样性监测能力,尤其适用于罕见及未知节肢动物的表征。普通视觉-语言模型(VLM)在常见物种分类中表现优异,而RAG模型通过匹配显式分类特征文本与外部生物多样性语料库,实现对稀有类群的准确分类。结果表明,RAG模型相比基础LLM降低了过度自信,提升了准确性,尤其在家族和属级分类中表现出更强的层次结构理解能力。研究强调了高质量数据整理与公民科学平台协作的重要性,为物种识别、未知物种分析及保护策略制定提供技术支持。
原文摘要 · Abstract (English)
In the context of pressing climate change challenges and the significant biodiversity loss among arthropods, automated taxonomic classification from organismal images is a subject of intense research. However, traditional AI pipelines based on deep neural visual architectures such as CNNs or ViTs face limitations such as degraded performance on the long-tail of classes and the inability to reason about their predictions. We integrate image captioning and retrieval-augmented generation (RAG) with large language models (LLMs) to enhance biodiversity monitoring, showing particular promise for characterizing rare and unknown arthropod species. While a naive Vision-Language Model (VLM) excels in classifying images of common species, the RAG model enables classification of rarer taxa by matching explicit textual descriptions of taxonomic features to contextual biodiversity text data from external sources. The RAG model shows promise in reducing overconfidence and enhancing accuracy relative to naive LLMs, suggesting its viability in capturing the nuances of taxonomic hierarchy, particularly at the challenging family and genus levels. Our findings highlight the potential for modern vision-language AI pipelines to support biodiversity conservation initiatives, emphasizing the role of comprehensive data curation and collaboration with citizen science platforms to improve species identification, unknown species characterization and ultimately inform conservation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。