arXiv:2505.24492cs.LGcs.AI2025-05NeurIPS被引 12

将物体中心模型与概念模型结合,提升视觉任务的性能与可解释性。

Object Centric Concept Bottlenecks

  • 用物体级编码替代全局图像编码,增强对复杂场景的理解能力。
  • 在复杂数据集上表现优于传统概念模型,支持多对象推理。
  • 适合需要透明决策的高阶视觉应用,如医疗影像分析。

构建高性能且可解释的模型仍是现代人工智能的关键挑战。概念基础模型(CBMs)通过从全局编码(如图像编码)中提取人类可理解的概念,并在概念激活上使用线性分类器,实现透明决策。然而,其依赖整体图像编码,在物体中心的真实场景中表达能力受限,难以解决单标签分类以外的复杂视觉任务。为此,我们提出物体中心概念瓶颈(OCB)框架,融合CBM与预训练物体中心基础模型的优势,提升性能与可解释性。我们在复杂图像数据集上评估OCB,通过全面消融实验分析关键组件,如物体-概念编码的聚合策略。结果表明,OCB优于传统CBMs,能为复杂视觉任务提供可解释的决策。

原文摘要 · Abstract (English)

Developing high-performing, yet interpretable models remains a critical challenge in modern AI. Concept-based models (CBMs) attempt to address this by extracting human-understandable concepts from a global encoding (e.g., image encoding) and then applying a linear classifier on the resulting concept activations, enabling transparent decision-making. However, their reliance on holistic image encodings limits their expressiveness in object-centric real-world settings and thus hinders their ability to solve complex vision tasks beyond single-label classification. To tackle these challenges, we introduce Object-Centric Concept Bottlenecks (OCB), a framework that combines the strengths of CBMs and pre-trained object-centric foundation models, boosting performance and interpretability. We evaluate OCB on complex image datasets and conduct a comprehensive ablation study to analyze key components of the framework, such as strategies for aggregating object-concept encodings. The results show that OCB outperforms traditional CBMs and allows one to make interpretable decisions for complex visual tasks.

可解释性物体中心概念模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。