动态调整噪声以提升多模态表示学习的鲁棒性
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
- 根据特征分布动态调节噪声强度,替代固定噪声
- 在多个视觉语言模型上实现一致性能提升
- 适合追求模型泛化能力的研究者与工程师
表示学习是现代机器学习的核心,支撑文本检索与多模态理解等应用。然而,学习鲁棒且可泛化的表示仍具挑战。尽管已有研究证明主动注入噪声(一种数据增强)能提升编码性能,但多数方法依赖启发式或静态噪声,忽视了训练过程中特征分布的动态变化。本文从梯度和特征分布两个角度系统研究噪声在表示学习中的作用,以InfoNCE损失为例。针对多模态表示学习,提出FANoise——一种基于特征自适应的噪声注入策略。通过利用对比学习的动态特性,FANoise有效缓解噪声的负面影响,同时保留其优势。在多种基础视觉语言模型上的实验表明,该方法在多模态任务中持续提升整体性能。
原文摘要 · Abstract (English)
Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generalizable representations remains challenging. While prior work has demonstrated that active noise injection, a form of data augmentation, can enhance encoding performance, most existing methods rely on heuristic or static noise, overlooking the dynamic nature of feature distributions during training. In this work, we systematically study the role of noise in representation learning from both gradient-based and feature distribution perspectives, using InfoNCE loss as a representative example. Focusing on multimodal representation learning, we propose FANoise, a novel feature-adaptive noise injection strategy. By leveraging the dynamics of contrastive learning, FANoise effectively mitigates the negative impacts of noise while preserving its benefits. Under this theoretically grounded framework, comprehensive experiments demonstrate that FANoise consistently improves overall performance on multimodal tasks across various base VLM models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。