用量子方法提升多模态语言理解,让机器更懂语法结构。
Multimodal Quantum Natural Language Processing: A Novel Framework for using Quantum Methods to Analyse Real Data
- 用量子计算框架分析语言语法,融合图像与文本数据
- 语法模型在图文分类中表现优于词袋和序列模型
- 适合关注量子自然语言处理的科研人员
尽管量子计算在多个领域取得进展,但将其应用于语言组合性建模(如语法结构与语义交互)的研究仍有限,尤其在整合图像、视频、音频等真实世界数据方面。本文探索量子计算方法如何通过多模态数据融合增强语言组合建模能力,提出并推进多模态量子自然语言处理(MQNLP)。利用Lambeq工具包,对比分析四种组合模型在图像-文本分类任务中的表现。结果显示,基于语法的模型(如DisCoCat和TreeReader)能更有效捕捉语法结构,而词袋和序列模型因缺乏语法感知能力表现较差。这些发现表明,量子方法在语言建模中具有潜力,未来随着量子技术发展或带来突破。
原文摘要 · Abstract (English)
Despite significant advances in quantum computing across various domains, research on applying quantum approaches to language compositionality - such as modeling linguistic structures and interactions - remains limited. This gap extends to the integration of quantum language data with real-world data from sources like images, video, and audio. This thesis explores how quantum computational methods can enhance the compositional modeling of language through multimodal data integration. Specifically, it advances Multimodal Quantum Natural Language Processing (MQNLP) by applying the Lambeq toolkit to conduct a comparative analysis of four compositional models and evaluate their influence on image-text classification tasks. Results indicate that syntax-based models, particularly DisCoCat and TreeReader, excel in effectively capturing grammatical structures, while bag-of-words and sequential models struggle due to limited syntactic awareness. These findings underscore the potential of quantum methods to enhance language modeling and drive breakthroughs as quantum technology evolves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。