为大模型可解释性提供统一评估框架,量化五种方法优劣。
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
- 构建四维新指标体系,系统评估XAI技术效果。
- LIME在多模型上表现稳定,AMV鲁棒性强且一致性近乎完美。
- 适合关注大模型可信度与XAI选型的研究者参考。
大语言模型(LLMs)日益复杂,对其透明性和可解释性的挑战凸显,亟需可解释人工智能(XAI)技术以增强可信度和可用性。本研究提出一个综合评估框架,包含四项新指标,用于衡量五种XAI技术在五种大模型及两个下游任务中的有效性。我们采用IMDB影评和推文情感抽取数据集,评估LIME、SHAP、集成梯度、层间相关性传播(LRP)和注意力机制可视化(AMV)等XAI方法。评估聚焦于四个关键指标:人类推理一致性(HA)、鲁棒性、一致性和对比性。结果表明,LIME在多种模型和指标下均表现优异;AMV展现出更强的鲁棒性与近乎完美的持续性;LRP在复杂模型中对比性表现突出。研究揭示了不同XAI方法的优缺点,为大模型中XAI技术的选择与开发提供重要指导。
原文摘要 · Abstract (English)
The increasing complexity of LLMs presents significant challenges to their transparency and interpretability, necessitating the use of eXplainable AI (XAI) techniques to enhance trustworthiness and usability. This study introduces a comprehensive evaluation framework with four novel metrics for assessing the effectiveness of five XAI techniques across five LLMs and two downstream tasks. We apply this framework to evaluate several XAI techniques LIME, SHAP, Integrated Gradients, Layer-wise Relevance Propagation (LRP), and Attention Mechanism Visualization (AMV) using the IMDB Movie Reviews and Tweet Sentiment Extraction datasets. The evaluation focuses on four key metrics: Human-reasoning Agreement (HA), Robustness, Consistency, and Contrastivity. Our results show that LIME consistently achieves high scores across multiple LLMs and evaluation metrics, while AMV demonstrates superior Robustness and near-perfect Consistency. LRP excels in Contrastivity, particularly with more complex models. Our findings provide valuable insights into the strengths and limitations of different XAI methods, offering guidance for developing and selecting appropriate XAI techniques for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。