用使用频率定义语言,重新评估大模型的语言建模能力。
Language Models Model Language
- 以使用频率为核心,将语言视为所有说写内容的总和。
- 反驳传统批判,认为大模型具备语言建模合理性。
- 适合语言学、自然语言处理研究者参考。
针对大语言模型(LLM)的语言学评论常受索绪尔与乔姆斯基理论框架影响,多为推测性且缺乏实效。批评者质疑模型是否能真正建模语言,强调需具备‘深层结构’或‘语义接地’才能达到理想化的语言‘能力’。本文主张转向经验主义哲学家维托德·马涅扎克的理论视角:语言并非‘符号系统’或‘大脑计算系统’,而是所有说写行为的总和,其核心驱动因素是语言元素的使用频率。基于此框架,我们挑战既有对大模型的批判,并提供设计、评估与解读语言模型的建设性路径。
原文摘要 · Abstract (English)
Linguistic commentary on LLMs, heavily influenced by the theoretical frameworks of de Saussure and Chomsky, is often speculative and unproductive. Critics challenge whether LLMs can legitimately model language, citing the need for "deep structure" or "grounding" to achieve an idealized linguistic "competence." We argue for a radical shift in perspective towards the empiricist principles of Witold Mańczak, a prominent general and historical linguist. He defines language not as a "system of signs" or a "computational system of the brain" but as the totality of all that is said and written. Above all, he identifies frequency of use of particular language elements as language's primary governing principle. Using his framework, we challenge prior critiques of LLMs and provide a constructive guide for designing, evaluating, and interpreting language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。