arXiv:2508.00545cs.LGcs.AI2025-08被引 3

提出可操作的可解释性定义,指导模型设计。

Foundations of Interpretable Models

  • 定义可解释性为通用、简洁且涵盖现有概念的框架
  • 揭示可解释模型所需的基础属性与架构特征
  • 开源首个支持可解释数据结构的工具库

我们指出,当前可解释性的定义不可操作,无法为用户提供建设性指导,导致可解释性研究本质存在问题。为此,我们提出一个通用、简洁且涵盖社区内非正式概念的可解释性定义。该定义具有可操作性,能直接揭示设计可解释模型所需的基础属性、底层假设、原则、数据结构和架构特征。基于此,我们提出一种通用模型设计蓝图,并发布首个原生支持可解释数据结构与流程的开源库。

原文摘要 · Abstract (English)

We argue that existing definitions of interpretability are not actionable in that they fail to inform users about general, sound, and robust interpretable model design. This makes current interpretability research fundamentally ill-posed. To address this issue, we propose a definition of interpretability that is general, simple, and subsumes existing informal notions within the interpretable AI community. We show that our definition is actionable, as it directly reveals the foundational properties, underlying assumptions, principles, data structures, and architectural features necessary for designing interpretable models. Building on this, we propose a general blueprint for designing interpretable models and introduce the first open-sourced library with native support for interpretable data structures and processes.

可解释性模型设计开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。