解决多模态学习中模态主导与信息冗余问题,提升预测可解释性。
Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation
- 双流结构分离模态特有与共享特征,通过残差投影实现解耦。
- 引入语义对齐头和正交约束,抑制跨模态冗余并防止特征坍塌。
- 在教育数据集上优于主流融合方法,适合需要可解释性的场景。
多模态学习已成为整合图像、文本和结构化属性等异构信息的基础范式。然而,多模态表示常受模态主导、信息冗余及虚假跨模态关联影响,导致泛化能力不足且可解释性差。高方差模态往往掩盖弱但语义重要的信号,而简单融合策略会无控制地纠缠共享与特有因子。为此,本文提出双流残差语义去相关网络(DSRSD-Net),通过残差分解与显式语义去相关约束,解耦模态特有与共享信息。该方法包含:(1) 双流表示学习模块,利用残差投影分离模态内(私有)与模态间(共享)潜在因子;(2) 残差语义对齐头,结合对比与回归目标将不同模态的共享因子映射至统一空间;(3) 去相关与正交性损失,规范共享空间协方差结构,并强制共享与私有流正交,从而抑制冗余并防止特征坍塌。在两个大规模教育基准上的实验表明,DSRSD-Net在下一步预测与最终结果预测任务上持续优于单模态、早期融合、晚期融合及协同注意力基线。
原文摘要 · Abstract (English)
Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dominance, redundant information coupling, and spurious cross-modal correlations, leading to suboptimal generalization and limited interpretability. In particular, high-variance modalities tend to overshadow weaker but semantically important signals, while naïve fusion strategies entangle modality-shared and modality-specific factors in an uncontrolled manner. This makes it difficult to understand which modality actually drives a prediction and to maintain robustness when some modalities are noisy or missing. To address these challenges, we propose a Dual-Stream Residual Semantic Decorrelation Network (DSRSD-Net), a simple yet effective framework that disentangles modality-specific and modality-shared information through residual decomposition and explicit semantic decorrelation constraints. DSRSD-Net introduces: (1) a dual-stream representation learning module that separates intra-modal (private) and inter-modal (shared) latent factors via residual projection; (2) a residual semantic alignment head that maps shared factors from different modalities into a common space using a combination of contrastive and regression-style objectives; and (3) a decorrelation and orthogonality loss that regularizes the covariance structure of the shared space while enforcing orthogonality between shared and private streams, thereby suppressing cross-modal redundancy and preventing feature collapse. Experimental results on two large-scale educational benchmarks demonstrate that DSRSD-Net consistently improves next-step prediction and final outcome prediction over strong single-modality, early-fusion, late-fusion, and co-attention baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。