用量子电路实现高效多模态融合,突破传统方法的参数瓶颈。
Expressive and Scalable Quantum Fusion for Multimodal Learning
- 用可调量子电路学习跨模态高阶交互,参数量线性增长
- 小规模任务上超越强基线,高模态场景下提升显著
- 适合追求可扩展多模态融合的量子机器学习研究者
本文提出一种量子融合层(Quantum Fusion Layer, QFL),用于多模态学习中的特征融合。与传统方法不同,QFL采用混合量子-经典架构,利用参数化量子电路学习模态间的纠缠特征交互,无需指数级参数增长。基于量子信号处理原理,其量子组件能以线性参数规模高效表示高阶多项式交互,并在模拟实验中展示了与低秩张量方法的分离性,体现潜在量子查询优势。在多个小型但多样化的多模态任务上,QFL持续优于主流经典基线,尤其在高模态场景下表现突出。结果表明,QFL提供了一种根本性且可扩展的多模态融合新范式,值得在更大系统上深入探索。
原文摘要 · Abstract (English)
The aim of this paper is to introduce a quantum fusion mechanism for multimodal learning and to establish its theoretical and empirical potential. The proposed method, called the Quantum Fusion Layer (QFL), replaces classical fusion schemes with a hybrid quantum-classical procedure that uses parameterized quantum circuits to learn entangled feature interactions without requiring exponential parameter growth. Supported by quantum signal processing principles, the quantum component efficiently represents high-order polynomial interactions across modalities with linear parameter scaling, and we provide a separation example between QFL and low-rank tensor-based methods that highlights potential quantum query advantages. In simulation, QFL consistently outperforms strong classical baselines on small but diverse multimodal tasks, with particularly marked improvements in high-modality regimes. These results suggest that QFL offers a fundamentally new and scalable approach to multimodal fusion that merits deeper exploration on larger systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。