解析AI决策逻辑,让黑箱模型可理解可信任
Explainable artificial intelligence (XAI): from inherent explainability to large language models
- 从可解释模型到大模型解释技术,系统梳理XAI发展脉络
- 利用大语言模型实现高阶语义级解释,提升决策透明度
- 适合关注AI可信性、安全性和落地应用的研究者与工程师
人工智能在近年取得巨大成功,但其决策过程往往不透明,导致利益相关方难以理解或解释其行为,阻碍了在医疗、自动驾驶等关键领域的信任与应用。可解释人工智能(XAI)技术通过增强机器学习模型的可解释性,使用户能够理解决策依据,防范不当行为。本文全面综述了从内在可解释模型到现代黑箱模型(包括大语言模型)的可解释性方法,还探讨了利用大语言模型和视觉-语言模型自动化或提升其他模型可解释性的技术。尤其值得注意的是,大模型能提供高层次、语义丰富的解释。文章剖析了前沿方法的科学原理、优缺点,并在适当处提供定性与定量对比,展示各方法性能差异。最后,讨论了当前面临的挑战及未来研究方向。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) has continued to achieve tremendous success in recent times. However, the decision logic of these frameworks is often not transparent, making it difficult for stakeholders to understand, interpret or explain their behavior. This limitation hinders trust in machine learning systems and causes a general reluctance towards their adoption in practical applications, particularly in mission-critical domains like healthcare and autonomous driving. Explainable AI (XAI) techniques facilitate the explainability or interpretability of machine learning models, enabling users to discern the basis of the decision and possibly avert undesirable behavior. This comprehensive survey details the advancements of explainable AI methods, from inherently interpretable models to modern approaches for achieving interpretability of various black box models, including large language models (LLMs). Additionally, we review explainable AI techniques that leverage LLM and vision-language model (VLM) frameworks to automate or improve the explainability of other machine learning models. The use of LLM and VLM as interpretability methods particularly enables high-level, semantically meaningful explanations of model decisions and behavior. Throughout the paper, we highlight the scientific principles, strengths and weaknesses of state-of-the-art methods and outline different areas of improvement. Where appropriate, we also present qualitative and quantitative comparison results of various methods to show how they compare. Finally, we discuss the key challenges of XAI and directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。