用量子电路增强多模态融合特征,参数少且更鲁棒。
PQFA: Parallel Quantum Feature Augmentation of Fused Representations for Multimodal Classification

- 用多个浅层量子电路对融合特征做并行增强。
- 在两个数据集上均优于经典方法,仅需2.2K参数。
- 对缺失模态有强鲁棒性,尤其文本受损时表现好。
多数多模态学习方法聚焦异构表征的对齐与融合,而融合后的增强研究较少。本文提出并行量子特征增强(PQFA),一种混合量子-经典框架,通过多个浅层变分量子电路对融合后的多模态特征进行处理。使用冻结的RoBERTa和ViT编码器提取文本与图像表示,经双向交叉注意力、注意力池化与自适应门控融合后,将融合特征进行振幅编码,输入并行量子电路,其测量结果与原始特征拼接用于预测。在MM-IMDb和N24News数据集上,使用相同编码器、融合主干、数据划分、投影维度和增强输出宽度进行受控对比。PQFA始终优于无量子增强的融合主干及宽度匹配的MLP基线,仅需约2.2K增强参数,远低于MLP分支的24.0K。缺失模态实验显示,当文本或视觉输入不完整时性能更稳定,尤其在文本模态严重退化时提升显著。受控消融与特征空间分析表明,性能提升无法由随机映射、增大经典宽度或未训练量子变换复制。量子态诊断进一步显示,在测试的模拟噪声水平下性能稳定,且各分支对编码态有特异性变换。结果证明PQFA是高效且参数节约的混合量子-经典多模态学习后融合增强策略。
原文摘要 · Abstract (English)
Most multimodal learning methods improve how heterogeneous representations are aligned and fused, while post-fusion enhancement remains less explored. We propose Parallel Quantum Feature Augmentation (PQFA), a hybrid quantum-classical framework that applies multiple shallow variational quantum circuits to fused multimodal features. Text and image representations extracted by frozen RoBERTa and ViT encoders are processed through bidirectional cross-attention, attentive pooling, and adaptive gated fusion. The fused feature is then amplitude-encoded into parallel quantum circuits, whose measurement readouts are concatenated with the classical representation for prediction. We evaluate PQFA on MM-IMDb and N24News through controlled comparisons using the same encoders, fusion backbone, data splits, projection dimension, and augmentation output width. PQFA consistently outperforms both the fusion backbone without quantum augmentation and a width-matched MLP augmentation baseline, while using approximately 2.2K augmentation parameters compared with 24.0K for the MLP branch. Missing-modality experiments further show improved robustness when textual or visual inputs are incomplete, with particularly clear gains when the more informative textual modality is severely degraded. Controlled ablations and feature-space analyses indicate that the improvement cannot be reproduced by random feature mappings, increased classical width, or untrained quantum transformations. Quantum-state diagnostics additionally show stable predictive performance across the tested simulated noise levels and distinct branch-specific transformations of the encoded states. These results establish PQFA as an effective and parameter-efficient strategy for post-fusion augmentation in hybrid quantum-classical multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。