用复数希尔伯特空间建模语言,让意义随上下文动态生成。
Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space
- 基于复数嵌入与共轭内积,构建可动态积累语义的复值序列模型
- 在相同训练条件下,参数越多越快逼近低损失,100M参数时表现优于实数模型
- 或可用更少参数实现当前大模型能力,适合消费级硬件部署
人类与大模型在自然语言处理中的实验表明,语义表达的意义在解释前是不确定的,而非简单由成分相加决定(即非组合性)。这种依赖观察者的动态过程更符合量子逻辑机制,而非假设分离性的经典布尔逻辑。为此,本文提出复数希尔伯特空间下的相位关联记忆(PAM)——一种状态 $ S_t \in \mathbb{C}^{d \times d} $ 的复值序列模型,通过共轭内积 $ \mathrm{Re}\langle K \mid Q\rangle / \sqrt{d} $ 检索复数词元嵌入并累积外积。在相同条件下对 WikiText-103 进行 5M 至 100M 参数量级的训练,两者均稳定训练;尽管 PAM 初期损失更高,但其损失下降幂律指数为 $-0.15$,困惑度为 $-0.65$,优于实数对照组的 $-0.12$ 和 $-0.49$,差距随参数增加持续缩小。未来大规模探索可能表明,类似 PAM 的架构可在仅需约 1000 亿参数(相比当前 1 万亿级前沿)的情况下达到现有大模型的损失平台,实现同等能力且可在消费级硬件运行。
原文摘要 · Abstract (English)
Experiments probing natural language processing by both humans and LLMs suggest that the meaning of a semantic expression is indeterminate prior to the act of interpretation rather than being specifiable simply as the sum of its parts (i.e. compositionality). This observer-dependent act dynamically actualizes meaning under genuine contextuality more consistent with quantum logical mechanisms than with classical Boolean approaches that assume separability, motivating an approach to language modeling that utilizes a Hilbert space formalism. In this work, we introduce Phase-Associative Memory (PAM) -- a complex-valued sequence model whose state S_t \in \mathbb{C}^{d \times d} accumulates outer products of complex token embeddings retrieved through the conjugate inner product $\mathrm{Re}\langle K \mid Q\rangle / \sqrt{d}$ -- and evaluate it against a structurally matched real-valued ablation. Both architectures train stably across a 5M--100M parameter sweep on WikiText-103 under identical conditions; PAM sits at higher absolute loss at every measured scale but improves more rapidly with parameter count, with power-law exponents of $-0.15$ vs.\ $-0.12$ in loss and $-0.65$ vs.\ $-0.49$ in perplexity that narrow the gap between the two architectures monotonically. Further investigation of complex-valued sequence modeling at larger scales could reveal that the loss plateau characteristic of real-valued state-of-the-art language models (e.g. transformers) is reachable with PAM-style architectures with an order of magnitude fewer parameters than the current frontier ($\sim$1T), implying that similar capabilities are achievable at sizes runnable on consumer-grade hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。