大模型融合符号与分布式表示,揭示语言理解新机制
LLMs as a synthesis between symbolic and distributed approaches to language
- 架构支持连续与离散表征并存,灵活切换
- 实证发现语法知识以近离散方式编码于模型中
- 适合研究语言认知与模型可解释性的人士阅读
自20世纪中期以来,符号主义与分布式方法在语言与认知领域长期对立。深度学习模型尤其是大语言模型的成功,或被视为分布式方法的胜利,或被贬低为工程进展。本文认为,语言类深度学习模型实际上实现了两种范式的融合:1)其架构同时支持分布式/连续/模糊与符号式/离散/分类表征与处理;2)训练后的语言模型利用这种灵活性。特别是,近期可解释性研究揭示,大量形态句法知识以近离散形式编码于大模型中。这表明不同行为以涌现方式产生,模型能根据需要在两种模式间灵活切换(及中间态)。这或许是其成功的关键原因之一,也使其成为语言研究的重要对象。是时候握手言和了吗?
原文摘要 · Abstract (English)
Since the middle of the 20th century, a fierce battle is being fought between symbolic and distributed approaches to language and cognition. The success of deep learning models, and LLMs in particular, has been alternatively taken as showing that the distributed camp has won, or dismissed as an irrelevant engineering development. In this position paper, I argue that deep learning models for language actually represent a synthesis between the two traditions. This is because 1) deep learning architectures allow for both distributed/continuous/fuzzy and symbolic/discrete/categorical-like representations and processing; 2) models trained on language make use of this flexibility. In particular, I review recent research in interpretability that showcases how a substantial part of morphosyntactic knowledge is encoded in a near-discrete fashion in LLMs. This line of research suggests that different behaviors arise in an emergent fashion, and models flexibly alternate between the two modes (and everything in between) as needed. This is possibly one of the main reasons for their wild success; and it makes them particularly interesting for the study of language. Is it time for peace?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。