用量子纠缠融合多模态信息,兼顾准确率与可解释性。
Feature Entanglement-based Quantum Multimodal Fusion Neural Network
- 通过量子纠缠实现跨模态特征融合,提升可解释性。
- 参数量仅为经典模型的十几分之一,准确率相当。
- 适合追求高效、可解释多模态融合的研究者。
多模态学习旨在通过整合多源信息提升感知与决策能力。然而,传统深度学习方法在特征级融合的高精度与决策级融合的可解释性之间存在根本矛盾,且面临参数爆炸与复杂度高的挑战。本文在量子计算框架下探讨了准确性-可解释性-复杂度的权衡问题,提出一种基于特征纠缠的量子多模态融合神经网络。该模型包含三个核心组件:用于单模态处理的经典前馈模块、可解释的量子融合模块,以及用于深层特征提取的量子卷积神经网络(QCNN)。借助量子强表达能力,将多模态融合与后处理的复杂度降至线性,并保留决策级融合的可解释性。仿真结果表明,该模型在多模态图像数据集上达到与经典网络相当的分类准确率,同时参数量仅为后者的十几分之一,展现出显著的稳定性与性能优势。
原文摘要 · Abstract (English)
Multimodal learning aims to enhance perceptual and decision-making capabilities by integrating information from diverse sources. However, classical deep learning approaches face a critical trade-off between the high accuracy of black-box feature-level fusion and the interpretability of less outstanding decision-level fusion, alongside the challenges of parameter explosion and complexity. This paper discusses the accuracy-interpretablity-complexity dilemma under the quantum computation framework and propose a feature entanglement-based quantum multimodal fusion neural network. The model is composed of three core components: a classical feed-forward module for unimodal processing, an interpretable quantum fusion block, and a quantum convolutional neural network (QCNN) for deep feature extraction. By leveraging the strong expressive power of quantum, we have reduced the complexity of multimodal fusion and post-processing to linear, and the fusion process also possesses the interpretability of decision-level fusion. The simulation results demonstrate that our model achieves classification accuracy comparable to classical networks with dozens of times of parameters, exhibiting notable stability and performance across multimodal image datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。