用量子力学框架解释大模型语义表示,揭示词义的波动与干涉特性。
Semantic Wave Functions: Exploring Meaning in Large Language Models Through Quantum Formalism
- 将语言嵌入扩展为复数域,构建语义波函数模拟词义叠加态。
- 引入相位与幅度联合相似度,提升对细微语义差异的感知能力。
- 适合对理论建模、语义动态演化感兴趣的学者参考。
大型语言模型(LLMs)在高维向量空间中编码语义关系。本文探索了LLM嵌入空间与量子力学之间的类比,提出LLMs运行于一种量化的语义空间中,词语和短语可视为量子态。为捕捉复杂的语义干涉效应,我们把标准实值嵌入空间扩展到复数域,借鉴双缝实验思想。引入“语义波函数”以形式化这一量子衍生表示,并利用势能景观(如双阱势)建模语义模糊性。此外,我们提出一种包含幅度与相位信息的复值相似度度量,实现更敏感的语义表征比较。基于非线性薛定谔方程与规范场及墨西哥帽势能,我们构建路径积分形式化框架,用于建模LLM行为的动态演化。该跨学科方法为理解并可能操控LLM提供了新的理论框架,旨在推进人工与自然语言理解。
原文摘要 · Abstract (English)
Large Language Models (LLMs) encode semantic relationships in high-dimensional vector embeddings. This paper explores the analogy between LLM embedding spaces and quantum mechanics, positing that LLMs operate within a quantized semantic space where words and phrases behave as quantum states. To capture nuanced semantic interference effects, we extend the standard real-valued embedding space to the complex domain, drawing parallels to the double-slit experiment. We introduce a "semantic wave function" to formalize this quantum-derived representation and utilize potential landscapes, such as the double-well potential, to model semantic ambiguity. Furthermore, we propose a complex-valued similarity measure that incorporates both magnitude and phase information, enabling a more sensitive comparison of semantic representations. We develop a path integral formalism, based on a nonlinear Schrödinger equation with a gauge field and Mexican hat potential, to model the dynamic evolution of LLM behavior. This interdisciplinary approach offers a new theoretical framework for understanding and potentially manipulating LLMs, with the goal of advancing both artificial and natural language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。