arXiv:2603.00197cs.CVcs.AI2026-03被引 1

用语义归纳法解析CNN神经元,验证其在新数据集的通用性。

A Case Study on Concept Induction for Neuron-Level Interpretability in CNN

  • 基于概念归纳框架,为神经元分配可解释语义标签。
  • 在SUN2012上成功识别出37个有意义的视觉概念。
  • 适合关注模型可解释性的研究人员使用。

深度神经网络在医疗、自动驾驶和场景理解等领域取得显著进展,但其隐藏层神经元的内部语义仍不清晰。先前工作提出了基于概念归纳的神经元分析框架,并在ADE20K数据集上验证了有效性。本案例研究进一步检验该方法在更大规模的场景识别基准SUN2012上的泛化能力。采用相同流程,为神经元赋予可解释的语义标签,并通过网络获取的图像与统计测试进行验证。结果表明,该方法可有效迁移至SUN2012,证实其具备更广泛的适用性。

原文摘要 · Abstract (English)

Deep Neural Networks (DNNs) have advanced applications in domains such as healthcare, autonomous systems, and scene understanding, yet the internal semantics of their hidden neurons remain poorly understood. Prior work introduced a Concept Induction-based framework for hidden neuron analysis and demonstrated its effectiveness on the ADE20K dataset. In this case study, we investigate whether the approach generalizes by applying it to the SUN2012 dataset, a large-scale scene recognition benchmark. Using the same workflow, we assign interpretable semantic labels to neurons and validate them through web-sourced images and statistical testing. Our findings confirm that the method transfers to SUN2012, showing its broader applicability.

可解释性CNN概念归纳

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。