用分阶段量子模型提升多模态组合泛化能力,参数量少且效果优。
Quantum Models with Multi-Stage Training for Compositional Concept Generalization

- 分两阶段训练:先学物体表征,再固定物体只优化关系部分。
- 在CLEVR数据集上,新方法在分布外关系泛化上显著优于经典模型。
- 非线性量子编码增强结构分离,适合研究量子多模态学习的学者。
组合概念泛化(CoCoGen)是多模态学习的核心挑战,即在新情境中系统重组已知基本元素的能力。本文提出一种基于意义组合的模型,将名词与关系分离,并利用张量和变分量子电路进行训练。采用分阶段训练策略:首先在单物体图像-文本对上学习物体表征,随后冻结物体参数,仅优化关系组件。该设计在量子电路层面显式实现组合分解,确保关系作为稳定基元的变换被学习。模型在专为CoCoGen设计的CLEVR数据集上验证。文本使用名词的向量表示与关系的高阶张量表示,搭配多种ansatz;图像则采用CLIP生成的图像嵌入进行量子编码,对比幅度编码(保留原始几何)与角度编码(引入非线性特征变换)。结果表明,分阶段训练结合结构化编码显著提升分布外关系泛化性能,且可训练参数量比经典基线少一个数量级。性能提升源于表示与编码的交互作用,其中非线性量子编码增强了组合结构的可分性。这些发现表明,结构化量子表示与分阶段学习为多模态量子机器学习中的组合泛化提供了有效框架。
原文摘要 · Abstract (English)
Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning. In this work, we provide a solution using a compositional model of meaning that separates nouns from relations and uses tensors and variational quantum circuits to train them on data. This model enables us to employ a multi stage training paradigm, one that first learns object representations from single-object image-caption pairs, then subsequently transfers these to the relational stage where object parameters are frozen and optimisation is only applied to relational components. This design explicitly enforces compositional factorisation at the circuit, ensuring that relations are learned as transformations over stable primitives. The training paradigm is tested on the CLEVR dataset developed specificially for CoCoGen. For text, we work with vector representations of nouns and higher order tensor representations of relations using a set of different ansatz. For images, we work with quantum encodings of image embeddings dervied from Open AI's Vision Language tool CLIP and contrast amplitude encoding, which preserves the original embedding geometry, with angle encoding, which introduces nonlinear feature transformations. Our results show that multi-staged training combined with structured encodings significantly improves out of distribution relational generalisation, while using orders of magnitude fewer trainable parameters than classical baselines. We find that performance gains arise from the interaction between representation and encoding, with nonlinear quantum encodings enhancing the separability of compositional structure. These findings demonstrate that structured quantum representations and staged learning provide an effective framework for compositional generalisation in multimodal quantum machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。