用范畴论构建逻辑一致的模型解释,提升可解释性可靠性。
Logic Explanation of AI Classifiers by Categorical Explaining Functors
- 基于范畴论设计解释函子,保持逻辑推理一致性
- 在合成数据上验证,矛盾与不忠实解释显著减少
- 适合需要严谨逻辑解释的AI可信场景
当前可解释人工智能主流方法为后验技术,通过识别预训练黑箱模型的关键特征来生成解释。部分先进方法能以逻辑规则形式捕捉输入特征间的相互作用,但常无法保证解释与模型底层推理的一致性。为此,本文提出一种理论基础扎实的方法,确保解释的连贯性与忠实度,突破现有启发式方法的局限。基于范畴论,引入解释函子,结构化保留解释与模型推理间的逻辑蕴含关系。作为概念验证,在合成基准上测试表明,该方法显著减少了矛盾或不忠实解释的生成。
原文摘要 · Abstract (English)
The most common methods in explainable artificial intelligence are post-hoc techniques which identify the most relevant features used by pretrained opaque models. Some of the most advanced post hoc methods can generate explanations that account for the mutual interactions of input features in the form of logic rules. However, these methods frequently fail to guarantee the consistency of the extracted explanations with the model's underlying reasoning. To bridge this gap, we propose a theoretically grounded approach to ensure coherence and fidelity of the extracted explanations, moving beyond the limitations of current heuristic-based approaches. To this end, drawing from category theory, we introduce an explaining functor which structurally preserves logical entailment between the explanation and the opaque model's reasoning. As a proof of concept, we validate the proposed theoretical constructions on a synthetic benchmark verifying how the proposed approach significantly mitigates the generation of contradictory or unfaithful explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。