提出多理由可解释图像识别新框架,提升判断准确性和理由质量。
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- 设计对比条件推理框架,显式建模图像、类别与理由的概率关系。
- 在多理由数据集上实现最佳零样本性能,分类与理由质量双优。
- 构建新基准数据集,支持更全面的可解释性评估,适合研究者使用。
基于视觉-语言模型(如CLIP)的可解释物体识别需输出准确类别标签,并提供支持决策的合理理由。现有方法多依赖提示条件,受限于CLIP文本编码器,且对解释结构的约束较弱。此外,先前数据集通常仅含单一、常有噪声的理由,无法充分捕捉判别性图像特征的多样性。本文提出一个包含多条真实理由标注的多理由可解释物体识别基准数据集,并设计相应评估指标以更全面反映任务特性。为克服前述局限,我们提出对比条件推理(CCI)框架,显式建模图像嵌入、类别标签与理由间的概率关系。该框架无需训练即可有效利用理由进行条件化,从而预测准确类别。在多理由可解释识别基准上,我们的方法达到最先进水平,包括优异的零样本表现,同时在分类准确率和理由质量上树立了新标准。结合该基准,本工作为未来可解释物体识别模型的评估提供了更完整的框架。代码将公开。
原文摘要 · Abstract (English)
Explainable object recognition using vision-language models such as CLIP involves predicting accurate category labels supported by rationales that justify the decision-making process. Existing methods typically rely on prompt-based conditioning, which suffers from limitations in CLIP's text encoder and provides weak conditioning on explanatory structures. Additionally, prior datasets are often restricted to single, and frequently noisy, rationales that fail to capture the full diversity of discriminative image features. In this work, we introduce a multi-rationale explainable object recognition benchmark comprising datasets in which each image is annotated with multiple ground-truth rationales, along with evaluation metrics designed to offer a more comprehensive representation of the task. To overcome the limitations of previous approaches, we propose a contrastive conditional inference (CCI) framework that explicitly models the probabilistic relationships among image embeddings, category labels, and rationales. Without requiring any training, our framework enables more effective conditioning on rationales to predict accurate object categories. Our approach achieves state-of-the-art results on the multi-rationale explainable object recognition benchmark, including strong zero-shot performance, and sets a new standard for both classification accuracy and rationale quality. Together with the benchmark, this work provides a more complete framework for evaluating future models in explainable object recognition. The code will be made available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。