用量子电路学习组合语义,提升图像描述的泛化能力。
Compositional Concept Generalization with Variational Quantum Circuits
- 将组合张量模型映射到希尔伯特空间,用变分量子电路学习表示
- 在噪声多热编码下实现概念组合泛化,效果优于经典模型
- 适合对量子机器学习与认知建模感兴趣的读者
组合泛化是人类认知的关键特征,但当前视觉-语言模型普遍缺乏此能力。以往研究尝试使用组合张量语义表示,但效果不佳。我们推测量子模型更高的训练效率可能改善此类任务表现。本文将组合张量模型的表示解释为希尔伯特空间中的向量,并训练变分量子电路,在需组合泛化的图像描述任务中学习这些表示。采用两种图像编码方式:对二值图像向量使用多热编码(MHE),对来自视觉-语言模型CLIP的图像向量使用角度/幅度编码。在带有噪声的MHE编码上取得良好概念验证结果;在CLIP图像向量上的表现较为波动,但仍优于经典组合模型。
原文摘要 · Abstract (English)
Compositional generalization is a key facet of human cognition, but lacking in current AI tools such as vision-language models. Previous work examined whether a compositional tensor-based sentence semantics can overcome the challenge, but led to negative results. We conjecture that the increased training efficiency of quantum models will improve performance in these tasks. We interpret the representations of compositional tensor-based models in Hilbert spaces and train Variational Quantum Circuits to learn these representations on an image captioning task requiring compositional generalization. We used two image encoding techniques: a multi-hot encoding (MHE) on binary image vectors and an angle/amplitude encoding on image vectors taken from the vision-language model CLIP. We achieve good proof-of-concept results using noisy MHE encodings. Performance on CLIP image vectors was more mixed, but still outperformed classical compositional models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。