arXiv:2602.09066cs.LGcs.AI2026-02被引 1

通过解耦特征频谱结构,提升多模态表示的鲁棒性与泛化能力。

Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning

  • 用SVD自适应划分特征维度为信号、噪声和冗余相关
  • 设计课程式增强策略,提升关键语义成分并保证训练稳定
  • 在特征与频谱双空间优化,适合追求高鲁棒性的多模态研究者

大规模多模态对比学习虽在获取丰富可迁移表示方面取得显著进展,但仍受限于对特征维度的均一处理及对学习特征内在频谱结构的忽视。实证表明,高维嵌入常坍缩为狭窄锥体,任务相关语义集中于小子空间,其余维度充斥噪声与虚假相关。这种频谱失衡与纠缠损害模型泛化能力。本文提出频谱解耦与增强(SDE)框架,连接嵌入空间几何与频谱特性。利用奇异值分解自适应划分特征维度为强信号(关键语义)、弱信号(辅助相关)和噪声(无关扰动)。采用课程式频谱增强策略,选择性放大信息成分,并具有训练稳定性理论保障。在此基础上,引入双域对比损失,联合优化特征空间与频谱空间的一致性,将频谱正则化融入训练过程,促进更丰富、更鲁棒的表示。大规模多模态基准实验表明,SDE持续提升表示鲁棒性与泛化性能,优于现有先进方法。SDE可无缝集成至现有对比学习流程,为多模态表示学习提供有效解决方案。

原文摘要 · Abstract (English)

Large-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the uniform treatment of feature dimensions and the neglect of the intrinsic spectral structure of the learned features. Empirical evidence indicates that high-dimensional embeddings tend to collapse into narrow cones, concentrating task-relevant semantics in a small subspace, while the majority of dimensions remain occupied by noise and spurious correlations. Such spectral imbalance and entanglement undermine model generalization. We propose Spectral Disentanglement and Enhancement (SDE), a novel framework that bridges the gap between the geometry of the embedded spaces and their spectral properties. Our approach leverages singular value decomposition to adaptively partition feature dimensions into strong signals that capture task-critical semantics, weak signals that reflect ancillary correlations, and noise representing irrelevant perturbations. A curriculum-based spectral enhancement strategy is then applied, selectively amplifying informative components with theoretical guarantees on training stability. Building upon the enhanced features, we further introduce a dual-domain contrastive loss that jointly optimizes alignment in both the feature and spectral spaces, effectively integrating spectral regularization into the training process and encouraging richer, more robust representations. Extensive experiments on large-scale multimodal benchmarks demonstrate that SDE consistently improves representation robustness and generalization, outperforming state-of-the-art methods. SDE integrates seamlessly with existing contrastive pipelines, offering an effective solution for multimodal representation learning.

多模态学习对比学习频谱分析表示增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。