从解释方法转向可解释模型,评估模型真实可理解性。
From Interpretability Methods to Interpretable Models
- 用现有工具分析不同模型的表征与计算内容
- 发现模型可解释性需由独立人类评估,而非专家确认
- 呼吁以模型为中心的可解释性研究新范式
十多年来,计算机视觉领域的可解释人工智能(XAI)已形成成熟的工具箱:归因、特征可视化、概念基和电路基方法。然而,大部分研究精力集中于构建和比较这些方法,而忽视了它们本应回答的核心问题——我们的模型究竟有多可解释?随着模型演进,我们是否在进步?本文主张将研究重点从方法转向模型,沿着两条互补路径推进:其一是已有工具足以描述和比较不同模型所表示与计算的内容;其二是更困难且被忽视的挑战——模型能否真正被依赖它的普通人理解,即独立评估者对模型的信任与认证所依赖的人类理解能力,这只能通过实证测量获得,无法推断。本文回顾工具成熟度,梳理模型对比研究的薄弱现状,类比系统神经科学,并提出以模型为中心的XAI新议程。
原文摘要 · Abstract (English)
More than a decade in, explainable AI (XAI) for computer vision has assembled a mature toolbox: attribution, feature visualization, concept-based, and circuit-based methods. Yet almost all of the field's effort has gone into building and comparing these methods, and little into the question they were meant to answer---how interpretable are our models, and are we making progress as they evolve? We argue for shifting the field's focus from methods to models, along two complementary lines. One is already within reach: existing tools let us characterize and compare what different models represent and compute. The other is harder, and largely neglected: whether a model can actually be understood by the humans who rely on it---the independent evaluators on whom trust and certification depend, not the experts confirming what they already expect. It can only be measured, not inferred. We review why the toolbox is mature enough to support both, survey the thin body of work comparing models, draw a parallel to systems neuroscience, and close with a model-centric XAI agenda.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。