语言不该是模型内部表示,而应作为边界接口与共享代码本。
Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition

- 将语言视为共享代码本的压缩机制,而非内部表征。
- 多模态模型在冲突线索中权重分配仅达理想值的11%至82%。
- 适合研究语言与认知、模型可解释性及系统设计的学者。
语言模型以标记为计算单位:语言是其输入、输出,也日益成为内部表示。语言是否应占据所有这些位置,取决于语言对使用它的系统所产生的影响。人类是唯一拥有百年数据研究此问题的系统。本文回顾语言如何影响人类感知、大脑和思维,并据此反思多模态模型与语言模型的设计。我们始终将语言视为运行在共享代码本上的压缩器:词是索引,内容由接收者解码,社区维护代码本。在人类中,这种压缩可度量,习得代码本会重组感官,且思维可在失去语言后依然存在。随后通过六种视觉-语言模型和两种机器人策略的线索冲突实验,测量模型对不一致线索的响应规则。结果显示,幸存线索的权重按其可靠性排序,达到理想观察者斜率的11%至82%,且许多答案直接复制文本。一类策略选择完全丢弃冗余线索而非降权,另一类则保持权重,在线索冲突时失效;一个在每帧训练中标识任务的视觉线索从未被学习,因语言路径已能拟合数据。语言模型是当前最接近人类语言网络的模型,它们已进入人类语言社区,改变词汇频率,而对齐过程缩小了其概念多样性。最后提出七条面向基于标记系统的启示:语言应位于模型边界与共享代码本中,如大脑一般,而非作为内部表示;脱离代码本的代价是可审计性丧失。
原文摘要 · Abstract (English)
Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。