解析大模型内部机制,提出可解释性新方法
A Theoretical Survey on Foundation Models
- 基于机器学习理论分析模型泛化、表达能力与动态行为
- 实现对大模型推理、训练及伦理问题的全流程解读
- 适合关注模型可信性与未来研究方向的研究者
理解黑箱基础模型(FMs)的内在机制在人工智能及其应用中至关重要但极具挑战。过去十年,研究聚焦于其可解释性,发展出事后解释方法以阐明模型已作出的具体决策。然而,这些方法在忠实度和资源消耗方面存在局限。因此,亟需考虑一类符合准确、全面、启发式且资源轻量要求的可解释方法,以揭示基础模型的底层机制。本综述旨在回顾满足上述原则并成功应用于基础模型的可解释方法。这些方法根植于机器学习理论,涵盖泛化性能、表达能力与动态行为的分析,为模型从推理能力、训练动态到伦理影响的整个流程提供深入解释。基于这些解读,本文识别了基础模型未来研究的关键前沿方向。
原文摘要 · Abstract (English)
Understanding the inner mechanisms of black-box foundation models (FMs) is essential yet challenging in artificial intelligence and its applications. Over the last decade, the long-running focus has been on their explainability, leading to the development of post-hoc explainable methods to rationalize the specific decisions already made by black-box FMs. However, these explainable methods have certain limitations in terms of faithfulness and resource requirement. Consequently, a new class of interpretable methods should be considered to unveil the underlying mechanisms of FMs in an accurate, comprehensive, heuristic, and resource-light way. This survey aims to review those interpretable methods that comply with the aforementioned principles and have been successfully applied to FMs. These methods are deeply rooted in machine learning theory, covering the analysis of generalization performance, expressive capability, and dynamic behavior. They provide a thorough interpretation of the entire workflow of FMs, ranging from the inference capability and training dynamics to their ethical implications. Ultimately, drawing upon these interpretations, this review identifies the next frontier research directions for FMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。