用视觉语言模型自动优化视觉概念库,提升识别效果。
Self-Evolving Visual Concept Library using Vision-Language Critics
- 通过视觉语言模型充当批评者,迭代优化概念库
- 在零样本、少样本和微调任务中均表现优异
- 无需人工标注,可直接部署使用
我们研究了为视觉识别构建视觉概念库的问题。手动定义概念费时费力,而完全依赖大语言模型生成概念又可能导致概念区分度不足或忽略概念间的复杂交互。本文提出ESCHER方法,从概念库学习视角出发,利用视觉语言模型(VLM)作为批评者,迭代发现并改进视觉概念,包括考虑概念间的相互作用及其对下游分类器的影响。通过利用大语言模型的上下文学习能力及历史性能记录,ESCHER根据VLM反馈动态优化概念生成策略。最终,ESCHER无需任何人工标注,是一个全自动的即插即用框架。我们在零样本、少样本和微调视觉分类任务上实证验证了其有效性。据我们所知,这是首个将概念库学习应用于真实视觉任务的工作。
原文摘要 · Abstract (English)
We study the problem of building a visual concept library for visual recognition. Building effective visual concept libraries is challenging, as manual definition is labor-intensive, while relying solely on LLMs for concept generation can result in concepts that lack discriminative power or fail to account for the complex interactions between them. Our approach, ESCHER, takes a library learning perspective to iteratively discover and improve visual concepts. ESCHER uses a vision-language model (VLM) as a critic to iteratively refine the concept library, including accounting for interactions between concepts and how they affect downstream classifiers. By leveraging the in-context learning abilities of LLMs and the history of performance using various concepts, ESCHER dynamically improves its concept generation strategy based on the VLM critic's feedback. Finally, ESCHER does not require any human annotations, and is thus an automated plug-and-play framework. We empirically demonstrate the ability of ESCHER to learn a concept library for zero-shot, few-shot, and fine-tuning visual classification tasks. This work represents, to our knowledge, the first application of concept library learning to real-world visual tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。