用大模型修正小样本识别错误,提升物种图像识别准确率
Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- 用大模型对小样本专家模型的预测结果进行事后修正
- 在5个数据集上平均提升6.4个百分点准确率
- 无需训练,适配多种模型和场景,实用性强
视觉物种识别(VSR)是生态学、古植物学、进化生物学等科学领域中实现物种级分类的基础任务。通过机器学习自动化VSR可显著加速研究进程,但物种标注需深厚专业知识,导致大规模带标签数据难获取,因此少样本学习(FSL)成为主流范式。与此同时,大型多模态模型(LMMs)展现出卓越的零样本视觉识别能力,引发其能否替代FSL专家模型用于VSR的思考。我们系统比较了FSL专家模型与当前LMMs的表现,发现尽管采用先进提示策略,现有LMMs仍显著落后于FSL模型。然而我们发现,当给定一张图像及专家模型生成的候选物种列表时,LMMs往往能纠正专家模型误判的顶级预测。基于此,我们提出无需训练的后处理修正框架(POC),通过多模态提示策略,使专家模型在五个VSR基准上平均提升6.4%准确率。实验表明POC可泛化至多种FSL方法、视觉编码器与LMM,是一种高效实用的VSR解决方案。
原文摘要 · Abstract (English)
Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology, palynology, evolutionary biology, systematics, and phylogenetics. Automating VSR through machine learning can significantly accelerate these efforts. However, species-level annotation requires extensive domain expertise, making large-scale labeled datasets difficult to obtain. Consequently, few-shot learning (FSL) is a practical paradigm, where an expert model is trained using only a few labeled examples. Meanwhile, Large Multimodal Models (LMMs) have demonstrated unprecedented zero-shot visual recognition capabilities, raising the question of whether they can serve as an alternative to FSL expert models for VSR. We start this work with a systematic comparison between FSL expert models and LMMs, revealing that, despite advanced prompting strategies, contemporary LMMs significantly underperform FSL expert models. Interestingly, we find that LMMs possess a complementary strength: given an image and a shortlist of candidate species generated by an expert model, LMMs can often recover the correct label when the expert model's top prediction is incorrect. Motivated by this, we propose Post-hoc Correction (POC), a simple training-free framework that leverages an LMM to post-process an expert model's top predictions. We develop a multimodal prompting strategy to enable POC to improve FSL expert models by 6.4 accuracy points, averaged over five VSR benchmarks. We show that POC generalizes across diverse FSL methods, visual encoders, and LMMs, making it a practical and effective framework for VSR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。