arXiv:2502.06536stat.MLcs.LG2025-02NeurIPS被引 6

无需干预即可高效学习可解释概念,理论保证学习正确性。

Sample-efficient Learning of Concepts with Theoretical Guarantees: from Data to Concepts without Interventions

  • 用无监督因果表示学习提取潜在因果变量,再少量标注对齐可解释概念。
  • 在强相关概念下仍保持低混淆度,准确率优于其他CBM方法。
  • 适合需要可解释性且标注成本高的场景,如医疗图像分析。

机器学习在许多现实系统中至关重要,但黑箱AI系统的可解释性、可解释性和鲁棒性仍存疑。概念瓶颈模型(CBM)通过从高维数据(如图像)中学习可解释概念来预测标签,缓解部分问题。然而,概念间的虚假相关会导致学习到“错误”概念。现有缓解策略依赖强假设,如概念间统计独立或需大量干预与人工标注。本文提出一个框架,在无需任何干预的情况下,提供所学概念正确性的理论保证及所需标签数的上限。该框架利用因果表示学习(CRL)方法,无监督地从高维观测中学习潜在因果变量,并通过少量概念标签将这些变量对齐到可解释概念。我们提出了线性和非参数两种估计器,分别提供线性情况下的有限样本高概率结果和非参数估计器的渐近一致性。在合成数据和图像基准测试中评估表明,所学概念杂质更少,且在概念强相关情况下准确率常高于其他CBM方法。

原文摘要 · Abstract (English)

Machine learning is a vital part of many real-world systems, but several concerns remain about the lack of interpretability, explainability and robustness of black-box AI systems. Concept Bottleneck Models (CBM) address some of these challenges by learning interpretable concepts from high-dimensional data, e.g. images, which are used to predict labels. An important issue in CBMs are spurious correlation between concepts, which effectively lead to learning "wrong" concepts. Current mitigating strategies have strong assumptions, e.g., they assume that the concepts are statistically independent of each other, or require substantial interaction in terms of both interventions and labels provided by annotators. In this paper, we describe a framework that provides theoretical guarantees on the correctness of the learned concepts and on the number of required labels, without requiring any interventions. Our framework leverages causal representation learning (CRL) methods to learn latent causal variables from high-dimensional observations in a unsupervised way, and then learns to align these variables with interpretable concepts with few concept labels. We propose a linear and a non-parametric estimator for this mapping, providing a finite-sample high probability result in the linear case and an asymptotic consistency result for the non-parametric estimator. We evaluate our framework in synthetic and image benchmarks, showing that the learned concepts have less impurities and are often more accurate than other CBMs, even in settings with strong correlations between concepts.

概念瓶颈因果学习可解释性小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。