大模型能准确描述自身决策的内部逻辑,还能通过训练提升解释能力。
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
- 让大模型在复杂决策中自述其权重偏好,实现自我解释。
- 微调后模型能更精准报告自身决策依据,且效果可迁移至新任务。
- 适合关注模型可解释性、安全控制的研究者与开发者。
我们对大型语言模型(LLMs)为何以特定方式回应的理解仍有限。其神经网络难以解析,个体神经元与电路的功能也尚未明晰。但另一种理解路径是探索并发展模型自我解释的能力。本文证明:一、在特定决策场景下,LLMs可准确描述自身内部过程的量化特征;二、通过训练可进一步提升此能力。我们对GPT-4o和GPT-4o-mini进行微调,使其在多种复杂情境(如选公寓、贷款、度假等)中,依据随机生成的定量偏好(如自然光与安静程度的相对重要性)做出决策。结果表明,模型能准确报告这些学习到的属性权重。进一步微调后,模型解释能力显著提升。更重要的是,该训练具有泛化能力,可改善模型对未训练过的复杂决策的解释精度。这项工作为训练大模型全面、准确地报告自身内部过程迈出了关键一步,有望极大提升模型的可解释性、可控性与安全性。
原文摘要 · Abstract (English)
We have only limited understanding of how and why large language models (LLMs) respond in the ways that they do. Their neural networks have proven challenging to interpret, and we are only beginning to tease out the function of individual neurons and circuits within them. However, another path to understanding these systems is to investigate and develop their capacity to explain their own functioning. Here, we show that i) LLMs can accurately describe quantitative features of their own internal processes during certain kinds of decision-making and ii) that it is possible to improve these capabilities through training. To do so, we fine-tuned GPT-4o and GPT-4o-mini to make decisions in a wide variety of complex contexts (e.g., choosing between condos, loans, vacations, etc.) according to randomly-generated, quantitative preferences about how to weigh different attributes (e.g., the relative importance of natural light versus quiet surroundings for condos). We demonstrate that the LLMs can accurately report these preferences (i.e., the weights that they learned to give to different attributes during decision-making). Next, we demonstrate that these LLMs can be fine-tuned to explain their decision-making even more accurately. Finally, we demonstrate that this training generalizes: It improves the ability of the models to accurately explain how they make other complex decisions, not just decisions they have been fine-tuned to make. This work is a step towards training LLMs to accurately and broadly report on their own internal processes -- a possibility that would yield substantial benefits for interpretability, control, and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。