大语言模型部分依赖表征推理,而非单纯记忆。
Representation in large language models
- 从认知机制角度分析模型行为,主张存在表征驱动。
- 提出可操作方法探测模型内部表征结构。
- 适合关注模型可解释性与智能本质的研究者。
近期大语言模型(LLMs)在多种任务上的卓越表现,引发了科学与哲学层面的广泛探讨,试图解释其工作机制。然而,对基本理论问题的分歧导致僵局,乐观派与悲观派对模型工作方式持截然不同观点。突破僵局需就根本问题达成共识,本文旨在回答:LLM行为是否部分由类生物认知的表征信息处理驱动,还是完全由记忆与随机查表过程决定?这是关于模型算法本质的问题,答案对判断系统是否具备信念、意图、概念、知识与理解具有深远意义。本文主张,LLM行为部分由表征驱动,并提出一系列实用技术以探测和解释这些表征。该框架为未来关于语言模型及其后继系统的理论研究奠定基础。
原文摘要 · Abstract (English)
The extraordinary success of recent Large Language Models (LLMs) on a diverse array of tasks has led to an explosion of scientific and philosophical theorizing aimed at explaining how they do what they do. Unfortunately, disagreement over fundamental theoretical issues has led to stalemate, with entrenched camps of LLM optimists and pessimists often committed to very different views of how these systems work. Overcoming stalemate requires agreement on fundamental questions, and the goal of this paper is to address one such question, namely: is LLM behavior driven partly by representation-based information processing of the sort implicated in biological cognition, or is it driven entirely by processes of memorization and stochastic table look-up? This is a question about what kind of algorithm LLMs implement, and the answer carries serious implications for higher level questions about whether these systems have beliefs, intentions, concepts, knowledge, and understanding. I argue that LLM behavior is partially driven by representation-based information processing, and then I describe and defend a series of practical techniques for investigating these representations and developing explanations on their basis. The resulting account provides a groundwork for future theorizing about language models and their successors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。