解析大模型决策过程,提升医疗与自动驾驶中的可信解释能力
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- 聚焦Transformer模型的局部可解释性与机制可解释性研究
- 在医疗和自动驾驶场景中验证解释对信任的影响
- 梳理当前挑战,提出可信解释的未来方向
大型语言模型在自然语言处理的众多下游任务中表现出色,但其预测下一个词和生成内容的内在机制对人类而言通常难以理解。此外,这些模型常出现预测和推理错误,即幻觉现象。这凸显了深入理解并解释语言模型内部运作机制的紧迫性,以及如何生成可信输出。本文致力于研究基于Transformer的大模型的局部可解释性与机制可解释性,以增强对模型的信任。主要贡献包括:首先,综述现有局部可解释性与机制可解释性方法及研究洞察;其次,开展在医疗与自动驾驶两大关键领域中关于可解释性与推理的实验研究,并分析解释对接收者信任的影响;最后,总结当前未解决的问题,展望生成与人类对齐、可信的大型语言模型解释所面临的机遇、关键挑战与未来方向。
原文摘要 · Abstract (English)
Large language models have exhibited impressive performance across a broad range of downstream tasks in natural language processing. However, how a language model predicts the next token and generates content is not generally understandable by humans. Furthermore, these models often make errors in prediction and reasoning, known as hallucinations. These errors underscore the urgent need to better understand and interpret the intricate inner workings of language models and how they generate predictive outputs. Motivated by this gap, this paper investigates local explainability and mechanistic interpretability within Transformer-based large language models to foster trust in such models. In this regard, our paper aims to make three key contributions. First, we present a review of local explainability and mechanistic interpretability approaches and insights from relevant studies in the literature. Furthermore, we describe experimental studies on explainability and reasoning with large language models in two critical domains -- healthcare and autonomous driving -- and analyze the trust implications of such explanations for explanation receivers. Finally, we summarize current unaddressed issues in the evolving landscape of LLM explainability and outline the opportunities, critical challenges, and future directions toward generating human-aligned, trustworthy LLM explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。