arXiv:2501.05855cs.CL2025-01ACL被引 10

用大模型模拟解释效果,实现可扩展的可解释性评估

ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability

  • 用大语言模型充当模拟器,自动验证概念解释是否可复现模型输出
  • 在多个模型和数据集上验证,大模型能稳定给出解释方法的排名
  • 适合需要大规模、一致化可解释性评估的研究者使用

基于概念的解释通过将复杂模型计算映射为人类可理解的概念来工作。评估这类解释极为困难,不仅涉及所生成概念空间的质量,还包括所选概念向用户传达的有效性。现有评估指标往往只关注前者,忽略后者。本文提出一种基于自动化可模拟性的评估框架:通过模拟器根据解释内容预测模型输出的能力来衡量解释有效性。该方法兼顾概念空间及其解释的可理解性,实现端到端评估。人工模拟可解释性研究因规模限制难以开展,尤其在全面实证评估中。为此,我们利用大语言模型(LLMs)作为模拟器,近似完成评估,并报告多种分析以确保近似可靠性。该方法支持跨模型与数据集的可扩展、一致评估。我们通过该框架进行综合性实证评估,结果表明大语言模型能提供一致的解释方法排序。代码已公开于 https://github.com/AnonymousConSim/ConSim。

原文摘要 · Abstract (English)

Concept-based explanations work by mapping complex model computations to human-understandable concepts. Evaluating such explanations is very difficult, as it includes not only the quality of the induced space of possible concepts but also how effectively the chosen concepts are communicated to users. Existing evaluation metrics often focus solely on the former, neglecting the latter. We introduce an evaluation framework for measuring concept explanations via automated simulatability: a simulator's ability to predict the explained model's outputs based on the provided explanations. This approach accounts for both the concept space and its interpretation in an end-to-end evaluation. Human studies for simulatability are notoriously difficult to enact, particularly at the scale of a wide, comprehensive empirical evaluation (which is the subject of this work). We propose using large language models (LLMs) as simulators to approximate the evaluation and report various analyses to make such approximations reliable. Our method allows for scalable and consistent evaluation across various models and datasets. We report a comprehensive empirical evaluation using this framework and show that LLMs provide consistent rankings of explanation methods. Code available at https://github.com/AnonymousConSim/ConSim.

可解释性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。