提出新型张量电路网络,低复杂度下实现更强泛化与抗噪能力。
A Tensor Residual Circuit Neural Network Factorized with Matrix Product Operation
- 融合张量网络与量子电路特性,用复数域并行激活提升非线性
- 在多个数据集上准确率比当前最优模型高2%-3%,抗噪声攻击仍有效
- 适合追求高效稳健模型的工程应用,尤其对噪声环境敏感场景
降低神经网络复杂度的同时保持泛化能力和鲁棒性是实际应用中的挑战。现有方法如基于克罗内克积的量子启发式网络和采用矩阵乘积算子(MPO)分解的混合张量网络,其全连接层的泛化性能不及量子电路模型。本文提出一种新型张量电路神经网络(TCNN),结合张量网络与残差电路模型的优势,在低复杂度下实现强泛化与鲁棒性。复数域中电路的并行激活操作提升了特征学习的非线性与效率。由于特征信息同时存在于参数的实部与虚部,设计了信息融合层以整合这些特征,增强泛化能力。实验表明,TCNN在多个数据集上的平均准确率较当前最优模型高出2%-3%;在参数扰动攻击下,其他模型失效,而TCNN仍具显著学习能力,归因于其防止梯度爆炸的能力。同时,其可训练参数量和CPU运行时间与对比模型相当。消融实验证明激活操作、并行结构和信息融合层均具关键作用。
原文摘要 · Abstract (English)
It is challenging to reduce the complexity of neural networks while maintaining their generalization ability and robustness, especially for practical applications. Conventional solutions for this problem incorporate quantum-inspired neural networks with Kronecker products and hybrid tensor neural networks with MPO factorization and fully-connected layers. Nonetheless, the generalization power and robustness of the fully-connected layers are not as outstanding as circuit models in quantum computing. In this paper, we propose a novel tensor circuit neural network (TCNN) that takes advantage of the characteristics of tensor neural networks and residual circuit models to achieve generalization ability and robustness with low complexity. The proposed activation operation and parallelism of the circuit in complex number field improves its non-linearity and efficiency for feature learning. Moreover, since the feature information exists in the parameters in both the real and imaginary parts in TCNN, an information fusion layer is proposed for merging features stored in those parameters to enhance the generalization capability. Experimental results confirm that TCNN showcases more outstanding generalization and robustness with its average accuracies on various datasets 2\%-3\% higher than those of the state-of-the-art compared models. More significantly, while other models fail to learn features under noise parameter attacking, TCNN still showcases prominent learning capability owing to its ability to prevent gradient explosion. Furthermore, it is comparable to the compared models on the number of trainable parameters and the CPU running time. An ablation study also indicates the advantage of the activation operation, the parallelism architecture and the information fusion layer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。