让大模型主动教人理解自己,实现更高效的互智能力解读。
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
- 通过多轮对话让大模型构建用户心智模型并主动解释自身行为。
- 相比传统黑箱分析,交互式解读提升人类对模型的理解深度。
- 适合希望深入理解大模型思维的开发者与研究者使用。
大语言模型(LLM)时代带来了新的可解释性机遇——代理式可解释性:在多轮对话中,模型主动构建对用户的认知模型,并据此协助人类建立对模型自身的更好理解。这种互动模式突破了传统‘检视式’可解释性(如打开黑箱)的局限。具备教学意图的大模型,其成功标准取决于人类是否真正理解。尽管代理式可解释性在高风险场景下可能因交互性而牺牲完整性,但其能借助合作模型发现潜在超越人类的概念,助力人类构建对机器的深层认知。该方法面临评估难题,因其具有‘人类嵌入反馈环’特性(人类响应是算法核心部分),我们讨论了应对方案与替代目标。随着大模型在多数任务上接近人类水平,代理式可解释性的价值在于帮助人类掌握模型可能具备的超人级认知,而非持续落后于理解。
原文摘要 · Abstract (English)
The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively assists human understanding by developing and leveraging a mental model of the user, which in turn enables humans to develop better mental models of the LLM. Such conversation is a new capability that traditional `inspective' interpretability methods (opening the black-box) do not use. Having a language model that aims to teach and explain--beyond just knowing how to talk--is similar to a teacher whose goal is to teach well, understanding that their success will be measured by the student's comprehension. While agentic interpretability may trade off completeness for interactivity, making it less suitable for high-stakes safety situations with potentially deceptive models, it leverages a cooperative model to discover potentially superhuman concepts that can improve humans' mental model of machines. Agentic interpretability introduces challenges, particularly in evaluation, due to what we call `human-entangled-in-the-loop' nature (humans responses are integral part of the algorithm), making the design and evaluation difficult. We discuss possible solutions and proxy goals. As LLMs approach human parity in many tasks, agentic interpretability's promise is to help humans learn the potentially superhuman concepts of the LLMs, rather than see us fall increasingly far from understanding them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。