通过因果干预解析视觉模型决策,发现隐藏概念关系与系统性偏差。
OCCAM: Open-set Causal Concept explAnation and Ontology induction for black-box vision Models

- 开放集发现视觉概念,用文本引导分割定位并干预移除
- 在Broden和ImageNet-S上提升解释质量,揭示全局概念结构
- 适合研究模型偏见、可解释性及视觉概念演化的人看
深度图像分类器的决策解释仍具挑战性,尤其在无法访问模型内部的黑箱设置下。我们提出OCCAM框架,用于视觉模型中的开放集因果概念解释与本体归纳。OCCAM以开放集方式发现视觉概念,通过文本引导分割进行定位,并通过移除概念的物体级干预来测量分类置信度变化,从而估算每个概念的因果贡献。除了局部解释外,OCCAM还聚合整个数据集的干预证据,构建出反映分类器如何全局组织视觉概念的结构化概念本体。基于该本体推理可揭示概念间一致依赖关系,暴露潜在因果关联,并揭示系统性模型偏差。在Broden和ImageNet-S多个分类器上的实验表明,OCCAM在开放集黑箱场景中提升了解释质量,且比单图归因方法提供更丰富的全局洞察。
原文摘要 · Abstract (English)
Interpreting the decisions of deep image classifiers remains challenging, particularly in black-box settings where model internals are inaccessible. We introduce OCCAM, a framework for open-set causal concept explanation and ontology induction in vision models. OCCAM discovers visual concepts in an open-set manner, localizes them via text-guided segmentation, and performs object-level interventions by removing concepts to measure changes in class confidence, estimating each concept's causal contribution. Beyond local explanations, OCCAM aggregates interventional evidence across a dataset to induce a structured concept ontology that captures how classifiers globally organize visual concepts. Reasoning over this ontology reveals consistent dependencies between concepts, exposes latent causal relations, and uncovers systematic model biases. Experiments on Broden and ImageNet-S across multiple classifiers show that OCCAM improves explanation quality in open-set black-box settings while providing richer global insight than per-image attribution methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。