揭示概念模型中的推理捷径问题及其可识别性条件
Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic Lens
- 从神经符号视角分析概念模型的推理捷径机制
- 提出理论条件,确保概念与推理层可正确识别
- 实证发现现有方法常无法满足这些条件
概念模型是通过概念提取器将输入映射到高层概念,并通过推理层将这些概念转化为预测的神经网络。确保这些模块产生可解释的概念并在分布外场景下可靠运行至关重要,但实现这一目标的条件仍不明确。本文通过建立概念模型与推理捷径(RSs)之间的新联系,研究该问题:即使推理层固定且预先给定,模型仍可能通过学习低质量概念获得高准确率。我们拓展了推理捷径的定义至更复杂的概念模型框架,并推导出识别概念与推理层的理论条件。实验结果表明推理捷径影响显著,现有方法即便结合多种自然缓解策略,也常无法满足这些理论条件。
原文摘要 · Abstract (English)
Concept-based Models are neural networks that learn a concept extractor to map inputs to high-level concepts and an inference layer to translate these into predictions. Ensuring these modules produce interpretable concepts and behave reliably in out-of-distribution is crucial, yet the conditions for achieving this remain unclear. We study this problem by establishing a novel connection between Concept-based Models and reasoning shortcuts (RSs), a common issue where models achieve high accuracy by learning low-quality concepts, even when the inference layer is fixed and provided upfront. Specifically, we extend RSs to the more complex setting of Concept-based Models and derive theoretical conditions for identifying both the concepts and the inference layer. Our empirical results highlight the impact of RSs and show that existing methods, even combined with multiple natural mitigation strategies, often fail to meet these conditions in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。