arXiv:2501.08281cs.LG2025-01被引 2

让深度模型生成可读逻辑规则,突破解释力与扩展性瓶颈

NEUROLOGIC: From Neural Representations to Interpretable Logic Rules

  • 从任意层神经表示直接提取逻辑规则,避免逐层替换的高成本
  • 在情感分析任务中成功生成有意义的可解释规则,效果优于现有方法
  • 支持复杂逻辑结构并融合人类先验知识,适合高阶认知任务解释

基于规则的解释方法能为神经网络行为提供严谨且全局可解释的洞察。然而,现有方法多局限于小型全连接网络,依赖昂贵的逐层规则提取与替换过程,难以推广至Transformer等复杂架构。此外,这些方法生成的规则过于浅显,类似决策树,无法捕捉计算机视觉和自然语言处理等复杂领域中的高层次抽象。为此,我们提出NEUROLOGIC框架,可直接从深层神经网络中提取可解释的逻辑规则。与以往方法不同,NEUROLOGIC能在任意选定层上构建基于隐藏谓词的逻辑规则,无需逐层重构,具备更强的架构兼容性与可扩展性。同时,该框架支持更丰富的逻辑表达,并可结合人类先验知识将隐藏谓词映射回输入空间,显著提升可解释性。我们在基于Transformer的情感分析任务上验证了NEUROLOGIC,结果表明其能提取出有意义、可解释的逻辑规则,为现有方法难以拓展的任务提供了更深层次的洞察。

原文摘要 · Abstract (English)

Rule-based explanation methods offer rigorous and globally interpretable insights into neural network behavior. However, existing approaches are mostly limited to small fully connected networks and depend on costly layerwise rule extraction and substitution processes. These limitations hinder their generalization to more complex architectures such as Transformers. Moreover, existing methods produce shallow, decision-tree-like rules that fail to capture rich, high-level abstractions in complex domains like computer vision and natural language processing. To address these challenges, we propose NEUROLOGIC, a novel framework that extracts interpretable logical rules directly from deep neural networks. Unlike previous methods, NEUROLOGIC can construct logic rules over hidden predicates derived from neural representations at any chosen layer, in contrast to costly layerwise extraction and rewriting. This flexibility enables broader architectural compatibility and improved scalability. Furthermore, NEUROLOGIC supports richer logical constructs and can incorporate human prior knowledge to ground hidden predicates back to the input space, enhancing interpretability. We validate NEUROLOGIC on Transformer-based sentiment analysis, demonstrating its ability to extract meaningful, interpretable logic rules and provide deeper insights-tasks where existing methods struggle to scale.

模型解释逻辑规则Transformer可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。