神经模型如何理解看似冗余的或句式,关键在上下文绑定
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
- 用神经机制解释为何重复词在上下文中不显冗余
- 语言模型通过上下文绑定词汇,避免语义重复
- 适合对语义动态和大模型认知感兴趣的读者
像“她会去法国或西班牙,或者也许是德国或法国”这样的句子形式上看似冗余,但在“玛丽将参加法国或西班牙的哲学项目,或德国或法国的数学项目”这类上下文中却可以接受。尽管传统分析依赖符号逻辑,本文提出基于人工神经网络的解释。我们首先通过人类与大语言模型的新行为实验,证明这种非冗余性在多种语境中具有稳定性。接着发现,语言模型中冗余规避源于两种相互作用的机制:模型学会将上下文相关信息绑定到重复词汇上,而Transformer的归纳头则选择性关注这些由上下文激活的表征。该神经解释揭示了上下文敏感语义解读的内在机制,补充了既有符号分析。
原文摘要 · Abstract (English)
Sentences like "She will go to France or Spain, or perhaps to Germany or France." appear formally redundant, yet become acceptable in contexts such as "Mary will go to a philosophy program in France or Spain, or a mathematics program in Germany or France." While this phenomenon has typically been analyzed using symbolic formal representations, we aim to provide an account grounded in artificial neural mechanisms. We first present new behavioral evidence from humans and large language models demonstrating the robustness of this apparent non-redundancy across contexts. We then show that, in language models, redundancy avoidance arises from two interacting mechanisms: models learn to bind contextually relevant information to repeated lexical items, and Transformer induction heads selectively attend to these context-licensed representations. We argue that this neural explanation sheds light on the mechanisms underlying context-sensitive semantic interpretation, and that it complements existing symbolic analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。