arXiv:2504.00040cs.CLcs.AI2025-04被引 1

用量子方法建模句子歧义,通过概率化处理过程提升语义表达

Quantum Methods for Managing Ambiguity in Natural Language Processing

  • 将句法歧义建模为过程的概率分布,而非仅词义
  • 用密度矩阵表示句子意义,可统一处理多种语言任务
  • 适合对量子自然语言处理感兴趣的研究者

范畴组合分布语义(DisCoCat)框架利用量子理论的数学结构,以形式化图示表示自然语言含义。这些图示可对应张量网络与量子电路,并在量子自然语言处理(QNLP)中与密度矩阵关联。以往的QNLP使用密度矩阵将歧义词建模为基本词的概率分布(如'queen'可能指君主或棋子)。本文提出将语法歧义建模为过程的概率分布,句子意义由密度矩阵表示。我们展示了如何构建代表句子意义的量子电路概率分布,并说明该方法可推广文献中的多项任务。通过实验验证了该理论的有效性。

原文摘要 · Abstract (English)

The Categorical Compositional Distributional (DisCoCat) framework models meaning in natural language using the mathematical framework of quantum theory, expressed as formal diagrams. DisCoCat diagrams can be associated with tensor networks and quantum circuits. DisCoCat diagrams have been connected to density matrices in various contexts in Quantum Natural Language Processing (QNLP). Previous use of density matrices in QNLP entails modelling ambiguous words as probability distributions over more basic words (the word \texttt{queen}, e.g., might mean the reigning queen or the chess piece). In this article, we investigate using probability distributions over processes to account for syntactic ambiguity in sentences. The meanings of these sentences are represented by density matrices. We show how to create probability distributions on quantum circuits that represent the meanings of sentences and explain how this approach generalises tasks from the literature. We conduct an experiment to validate the proposed theory.

量子自然语言歧义建模密度矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。