用自生成理论补全整合主义语言观,解释语言如何持续开放与积累。
How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism
- 提出自生成机制,解释语言如何保持未来行动的开放性
- 揭示语言与非语言符号活动的连续性计算基础
- 构建历史整合残留物的结构理论,适合语言模型研究者
罗伊·哈里斯的整合主义语言学批判了将语言视为映射既定世界的编码传统,主张语言是面向未来共同行动的情境化二元活动。然而整合主义存在解释空白:未阐明符号维持前瞻性开放的结构机制,低估语言与非语言符号活动之间的连续性,也缺乏对过往整合积累档案的详细描述。本文认为,针对大语言模型行为提出的伊兰·巴伦霍尔茨自生成语言理论可精准填补这些空白,在不违背整合主义核心立场的前提下,提供:语言前瞻性开放的结构机制;语言与其它符号活动连续性的计算对应;对过往整合残留物结构及新参与者如何利用它的理论。该综合保留了情境化整合行为的本体优先性,同时补充了整合主义自身无法提供的解释内容。对自然语言处理与大语言模型设计的研究者而言,该论证为语言模型所依赖的统计结构的本质及其内在局限提供了原则性解释。
原文摘要 · Abstract (English)
Roy Harris's Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational approaches to language, arguing that language is not a code that maps onto a pre-given world but a situated, bipartite activity oriented toward prospective joint action. Yet Integrationism leaves certain explanatory gaps: it does not fully account for the structural mechanism by which signs sustain prospective openness, it undertheorises the continuity between linguistic and non-linguistic semiotic activity, and it offers no detailed account of the structural properties of the accumulated archive of past integrations. This paper argues that Elan Barenholtz's autogenerative theory of language, developed in response to the behaviour of Large Language Models (LLMs), can fill precisely these gaps, enriching Integrationism without undermining any of its core commitments. Specifically, the autogenerative account provides: a structural mechanism for the prospective openness that Harris identifies as central to bipartite communication; a computational correlate for Harris's thesis of semiotic continuity between language and other sign-making activity; and a theory of the archive: what the accumulated residue of past integrations looks like and how new participants draw upon it. The synthesis preserves Harris's ontological primacy of the situated integrative act while adding explanatory content that Integrationism itself does not supply. For practitioners and researchers in natural language processing and large language model design, the argument offers a principled account of what the statistical structure that LLMs so effectively exploit actually is, and of what it cannot, by its nature, provide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。