arXiv:2506.10077cs.CLcs.AI2025-06中稿 · submission to Quan…被引 7

提出量子语义框架,解释语言歧义的非经典特性。

A quantum semantic framework for natural language processing

  • 用柯尔莫哥洛夫复杂度分析语言歧义的指数级增长
  • 实验显示语言理解违反经典贝尔不等式(最大2.8)
  • 适合关注语言本质与认知计算的学者

语义退化是自然语言的根本属性,不仅限于多义性,更表现为表达越复杂,潜在解释呈组合爆炸。本文指出,这一特性使大语言模型等现代NLP系统面临根本局限,因其运行在语言本身之中。基于柯尔莫哥洛夫复杂度,我们证明随着表达复杂度上升,恢复其确切含义所需上下文信息量呈组合爆炸式增长,导致恢复唯一意图在计算上不可行。这表明语言形式本身具有内在意义的观点在概念上不充分。相反,我们认为意义是通过观察者依赖的解释行为动态实现的,其非确定性特征更适合用非经典量子逻辑描述。为验证该假设,我们使用多种大语言模型代理进行了语义贝尔不等式测试,平均CHSH期望值达1.2至2.8,部分结果(如2.3-2.4)显著违反经典边界(|S|≤2),表明语言在歧义情境下的理解可表现出非经典关联性,与人类认知实验结果一致。这些结果暗示,基于经典频率主义的自然语言分析方法必然存在信息损失。因此,我们主张采用贝叶斯式的重复采样方法,以更实际、适切地刻画语境中的语言意义。

原文摘要 · Abstract (English)

Semantic degeneracy represents a fundamental property of natural language that extends beyond simple polysemy to encompass the combinatorial explosion of potential interpretations that emerges as semantic expressions increase in complexity. In this work, we argue this property imposes fundamental limitations on Large Language Models (LLMs) and other modern NLP systems, precisely because they operate within natural language itself. Using Kolmogorov complexity, we demonstrate that as an expression's complexity grows, the amount of contextual information required to reliably resolve its ambiguity explodes combinatorially. The computational intractability of recovering a single intended meaning for complex or ambiguous text therefore suggests that the classical view that linguistic forms possess intrinsic meaning in and of themselves is conceptually inadequate. We argue instead that meaning is dynamically actualized through an observer-dependent interpretive act, a process whose non-deterministic nature is most appropriately described by a non-classical, quantum-like logic. To test this hypothesis, we conducted a semantic Bell inequality test using diverse LLM agents. Our experiments yielded average CHSH expectation values from 1.2 to 2.8, with several runs producing values (e.g., 2.3-2.4) in significant violation of the classical boundary ($|S|\leq2$), demonstrating that linguistic interpretation under ambiguity can exhibit non-classical contextuality, consistent with results from human cognition experiments. These results inherently imply that classical frequentist-based analytical approaches for natural language are necessarily lossy. Instead, we propose that Bayesian-style repeated sampling approaches can provide more practically useful and appropriate characterizations of linguistic meaning in context.

语义学量子计算认知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。