arXiv:2505.21657cs.CLcs.AI2025-05被引 4

用扰动法分析大模型每一步推理,让黑箱决策变透明。

Explaining Large Language Models with gSMILE

  • 通过控制输入扰动+距离度量,定位影响输出的关键词元。
  • 在GPT-3.5和Claude 2.1上验证,对齐人类判断且结果稳定。
  • 适合需要可信解释的医疗、金融等高风险场景使用。

大型语言模型(如GPT、LLaMA、Claude)在文本生成中表现卓越,但其决策过程不透明,限制了其在高风险应用中的信任与问责。我们提出gSMILE(生成式SMILE),一种模型无关、基于扰动的词元级可解释性框架。扩展SMILE方法,gSMILE利用受控提示扰动、Wasserstein距离度量和加权线性代理,识别对输出影响最大的输入词元。该过程可生成直观热图,可视化关键词元与推理路径。我们在主流模型(OpenAI的gpt-3.5-turbo-instruct、Meta的LLaMA 3.1 Instruct Turbo、Anthropic的Claude 2.1)上评估,采用归因保真度、一致性、稳定性、忠实度和准确率等指标。结果显示,gSMILE能提供可靠的人类对齐归因:Claude 2.1在注意力保真度上最优,GPT-3.5输出一致性最高。这些发现证明gSMILE能在模型性能与可解释性间取得平衡,推动更透明、可信的AI系统发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) such as GPT, LLaMA, and Claude achieve remarkable performance in text generation but remain opaque in their decision-making processes, limiting trust and accountability in high-stakes applications. We present gSMILE (generative SMILE), a model-agnostic, perturbation-based framework for token-level interpretability in LLMs. Extending the SMILE methodology, gSMILE uses controlled prompt perturbations, Wasserstein distance metrics, and weighted linear surrogates to identify input tokens with the most significant impact on the output. This process enables the generation of intuitive heatmaps that visually highlight influential tokens and reasoning paths. We evaluate gSMILE across leading LLMs (OpenAI's gpt-3.5-turbo-instruct, Meta's LLaMA 3.1 Instruct Turbo, and Anthropic's Claude 2.1) using attribution fidelity, attribution consistency, attribution stability, attribution faithfulness, and attribution accuracy as metrics. Results show that gSMILE delivers reliable human-aligned attributions, with Claude 2.1 excelling in attention fidelity and GPT-3.5 achieving the highest output consistency. These findings demonstrate gSMILE's ability to balance model performance and interpretability, enabling more transparent and trustworthy AI systems.

大模型解释可解释性归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。