arXiv:2410.03964cs.LGcs.AI2024-10EMNLP被引 4

用变分贝叶斯方法实现大模型的语义概念级解释

Variational Language Concepts for Interpreting Foundation Language Models

  • 提出变分语言概念框架,从词级解释升级到概念级解释
  • 在多个真实数据集上验证了对大模型预测的有效解释能力
  • 适合关注大模型可解释性与语义理解的研究者

基础语言模型(FLMs)如BERT及其变体在自然语言处理中取得了显著成果。目前,对FLMs的可解释性研究主要依赖自注意力层中的注意力权重,但这些权重仅提供词级别解释,无法捕捉更高层次的结构,因此可读性和直观性不足。为解决此问题,我们首先形式化定义了概念解释,随后提出一种变分贝叶斯框架——变分语言概念(VALC),以突破词级解释的局限,实现概念级解释。理论分析表明,VALC能找出最优的语言概念来解释FLM的预测结果。在多个真实数据集上的实证结果表明,该方法能够成功为FLMs提供概念级解释。

原文摘要 · Abstract (English)

Foundation Language Models (FLMs) such as BERT and its variants have achieved remarkable success in natural language processing. To date, the interpretability of FLMs has primarily relied on the attention weights in their self-attention layers. However, these attention weights only provide word-level interpretations, failing to capture higher-level structures, and are therefore lacking in readability and intuitiveness. To address this challenge, we first provide a formal definition of conceptual interpretation and then propose a variational Bayesian framework, dubbed VAriational Language Concept (VALC), to go beyond word-level interpretations and provide concept-level interpretations. Our theoretical analysis shows that our VALC finds the optimal language concepts to interpret FLM predictions. Empirical results on several real-world datasets show that our method can successfully provide conceptual interpretation for FLMs.

可解释性语言模型变分推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。