arXiv:2601.13798cs.CVcs.AI2026-01中稿 · ECCV被引 4

让视觉模型的决策过程可解释,通过空间定位的语义概念生成清晰说明。

CFM: Language-aligned Concept Foundation Model for Vision

  • 构建语言对齐的视觉概念基础模型,实现细粒度且空间精准的概念提取。
  • 在分类、分割和图文生成任务上性能媲美黑箱模型,同时提供高质量解释。
  • 通过概念共现关系优化命名与关联,适合需要透明化AI决策的场景。

语言对齐的视觉基础模型在多种下游任务中表现优异,但其内部表征不透明,难以解释决策过程。现有方法虽能分解为人类可理解的概念,但空间定位差且仅适用于图像分类任务。本文提出CFM,一种面向视觉的语言对齐概念基础模型,能够生成细粒度、可解释且空间定位准确的概念。当与具备强语义表征的基础模型结合时,可为任意下游任务提供解释。通过分析概念的局部共现依赖关系,定义概念间关联,从而改进概念命名并生成更丰富的解释。在基准数据集上,CFM在分类、分割和图文生成任务上的表现与黑箱模型相当,同时提供高质细粒度的概念解释。代码开源于https://github.com/kawi19/CFM,交互式可视化展示见https://concept-foundation-model.mpi-inf.mpg.de。

原文摘要 · Abstract (English)

Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision-making difficult. Recent work decompose these representations into human-interpretable concepts, but provide poor spatial grounding and are limited to image classification tasks. In this work, we propose CFM, a language-aligned concept foundation model for vision that provides fine-grained concepts, which are human-interpretable and spatially grounded in the input image. When paired with a foundation model with strong semantic representations, we get explanations for any of its downstream tasks. Examining local co-occurrence dependencies of concepts allows us to define concept relationships through which we improve concept naming and obtain richer explanations. On benchmark data, we show that CFM provides performance on classification, segmentation, and captioning that is competitive with opaque foundation models while providing fine-grained, high quality concept-based explanations. Code at https://github.com/kawi19/CFM. Interactive visualizations at https://concept-foundation-model.mpi-inf.mpg.de.

可解释AI概念模型视觉理解空间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。