arXiv:2603.10084cs.LGcs.AI2026-03中稿 · ICLR

从粗粒度标注中自动构建多层级概念体系,支持可解释的模型干预。

Digging Deeper: Learning Multi-Level Concept Hierarchies

  • 通过多级概念拆分,仅用顶层标注发现深层概念结构。
  • 在多个数据集上识别出训练时未显式标注的人类可理解概念。
  • 支持多层抽象干预,提升任务表现且保持高准确率。

尽管基于概念的模型通过人类可理解的概念解释预测结果,但通常依赖详尽标注,并将概念视为扁平独立。为克服此问题,近期工作引入层次化概念嵌入模型(HiCEMs)与概念拆分技术,仅使用粗粒度标注发现子概念。然而,两者均局限于浅层层次结构。本文提出多层级概念拆分(MLCS),可仅从顶层监督中发现多层级概念层次结构;并设计Deep-HiCEMs架构,有效表示这些发现的层次结构,支持在多层抽象层面进行测试时干预。在多个数据集上的实验表明,MLCS能发现训练中未显式出现的人类可理解概念,而Deep-HiCEMs在保持高精度的同时,支持测试阶段的概念干预,进一步提升任务性能。

原文摘要 · Abstract (English)

Although concept-based models promise interpretability by explaining predictions with human-understandable concepts, they typically rely on exhaustive annotations and treat concepts as flat and independent. To circumvent this, recent work has introduced Hierarchical Concept Embedding Models (HiCEMs) to explicitly model concept relationships, and Concept Splitting to discover sub-concepts using only coarse annotations. However, both HiCEMs and Concept Splitting are restricted to shallow hierarchies. We overcome this limitation with Multi-Level Concept Splitting (MLCS), which discovers multi-level concept hierarchies from only top-level supervision, and Deep-HiCEMs, an architecture that represents these discovered hierarchies and enables interventions at multiple levels of abstraction. Experiments across multiple datasets show that MLCS discovers human-interpretable concepts absent during training and that Deep-HiCEMs maintain high accuracy while supporting test-time concept interventions that can improve task performance.

概念学习层次结构可解释性概念干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。