不用训练,用大模型给图像区域自动标注语义概念,准确率超60%。
Low-cost concept-based localized explanations: How far can we get with training-free approaches?

- 零样本提示+嵌入相似度,无需训练即可定位并命名图像中的物体和部件
- 在多个数据集上达到62%-88%的物体级别命名准确率
- 提供可复现框架,适合快速开展低成本可解释性研究
概念驱动的可解释AI(C-XAI)旨在提供基于语义概念的人类可理解解释,但受限于细粒度概念标注的稀缺。我们评估中等规模多模态大模型(MLLMs,7B-32B)在严格零样本条件下,能否对目标框区域进行局部化概念命名。提出可复现的零样本概念命名评估协议:(i) 闭集、类别约束提示用于中等词汇量;(ii) Open-CoNa,一种基于嵌入相似度的大标签空间策略。在四个MLLM上实验显示,跨数据集表现一致,物体级精确匹配准确率达62%-88%,表明无需训练即可从局部区域生成概念标注具有巨大潜力。讨论了局限性和失败模式,并发布可复现框架以支持未来低成本C-XAI研究。
原文摘要 · Abstract (English)
Concept-based Explainable AI (C-XAI) seeks human-understandable explanations grounded in semantic concepts, yet validation is limited by the scarcity of fine-grained concept annotations. We evaluate whether mid-scale Multimodal Large Language Models (MLLMs) can perform localized concept naming under strict zero-shot conditions by assigning labels to bounding-box regions at both object and part levels. We propose a reproducible zero-shot evaluation protocol for Concept Naming (CoNa) with (i) closed-set, category-constrained prompting for moderate vocabularies and (ii) Open-CoNa, an embedding-similarity-based strategy for large label spaces. Experiments with four MLLMs (7B-32B) show consistent performance trends across datasets, reaching 62%-88% object-level exact-match accuracy, highlighting the potential of training-free concept annotation from localized regions. We discuss limitations and failure modes and release a reproducible framework to support future low-cost C-XAI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。