用人类联想方式提升视觉模型零样本识别能力
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
- 基于贝叶斯推理,将概念作为隐变量动态组合
- 在15个数据集上超越现有最先进方法
- 适合需要灵活识别新类别的实际应用场景
在零样本图像识别任务中,人类能通过组合已知简单概念灵活分类未见类别。然而,现有视觉语言模型虽借助大规模自然语言监督取得进展,仍因提示工程不佳及对目标类适应性差,在真实场景表现受限。为此,我们提出概念引导的人类式贝叶斯推理(CHBR)框架。该框架基于贝叶斯定理,将人类图像识别中的概念视为隐变量,通过先验分布与似然函数加权求和来建模任务。为解决无限概念空间下的不可计算问题,引入重要性采样算法,迭代调用大语言模型生成具有区分性的概念,强调类别间差异。此外,提出平均似然、置信度似然与测试时增强似然三种启发式策略,根据测试图像动态优化概念组合。在十五个数据集上的广泛评估表明,CHBR持续优于现有最先进的零样本泛化方法。
原文摘要 · Abstract (English)
In zero-shot image recognition tasks, humans demonstrate remarkable flexibility in classifying unseen categories by composing known simpler concepts. However, existing vision-language models (VLMs), despite achieving significant progress through large-scale natural language supervision, often underperform in real-world applications because of sub-optimal prompt engineering and the inability to adapt effectively to target classes. To address these issues, we propose a Concept-guided Human-like Bayesian Reasoning (CHBR) framework. Grounded in Bayes' theorem, CHBR models the concept used in human image recognition as latent variables and formulates this task by summing across potential concepts, weighted by a prior distribution and a likelihood function. To tackle the intractable computation over an infinite concept space, we introduce an importance sampling algorithm that iteratively prompts large language models (LLMs) to generate discriminative concepts, emphasizing inter-class differences. We further propose three heuristic approaches involving Average Likelihood, Confidence Likelihood, and Test Time Augmentation (TTA) Likelihood, which dynamically refine the combination of concepts based on the test image. Extensive evaluations across fifteen datasets demonstrate that CHBR consistently outperforms existing state-of-the-art zero-shot generalization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。