arXiv:2411.04242cs.LG2024-11被引 3

用量子计算处理文本图像结构,提升模型可解释性

Multimodal Structure-Aware Quantum Data Processing

  • 将语言语法与图像层级结构映射到量子电路中
  • 在主流图像分类任务上达到顶尖性能,且模型完全结构化
  • 适合关注量子机器学习与可解释性的研究者

尽管大语言模型推动了自然语言处理发展,但其'黑箱'特性阻碍了决策过程的理解。为解决此问题,研究者采用高阶张量构建结构化方法,能建模语言关系,但在经典计算机上因规模过大而训练困难。张量天然适配量子系统,通过将文本转译为变分量子电路可在量子计算机上实现训练。本文提出MultiQ-NLP框架,用于处理多模态文本+图像数据的结构感知处理。其中'结构'指语言中的句法语义关系及图像中视觉元素的层次组织。我们引入新型张量类型与类型同态,并设计新架构表示结构。在主流图像分类任务(SVO Probes)上,最优模型表现与现有最佳经典模型相当;且最优模型完全基于结构化设计。

原文摘要 · Abstract (English)

While large language models (LLMs) have advanced the field of natural language processing (NLP), their "black box" nature obscures their decision-making processes. To address this, researchers developed structured approaches using higher order tensors. These are able to model linguistic relations, but stall when training on classical computers due to their excessive size. Tensors are natural inhabitants of quantum systems and training on quantum computers provides a solution by translating text to variational quantum circuits. In this paper, we develop MultiQ-NLP: a framework for structure-aware data processing with multimodal text+image data. Here, "structure" refers to syntactic and grammatical relationships in language, as well as the hierarchical organization of visual elements in images. We enrich the translation with new types and type homomorphisms and develop novel architectures to represent structure. When tested on a main stream image classification task (SVO Probes), our best model showed a par performance with the state of the art classical models; moreover the best model was fully structured.

量子计算多模态结构化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。