arXiv:2503.03202cs.CV2025-03

低数据下动态调整对比损失,提升图文对齐效果

Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings

  • 根据对齐预测的不确定性动态调节损失权重
  • 在Flickr8k子集上检索准确率提升,优于固定权重基线
  • 对噪声扰动更鲁棒,适合小样本多模态任务

在低数据场景下,标准对比学习易因过拟合和训练不稳定导致模态对齐效果差。本文提出一种方差感知的损失调度方法,根据模型对齐预测的统计变异性(不确定性)动态调整对比损失权重。基于Flickr8k数据集子集模拟有限数据条件,实验表明该方法相比固定权重基线显著提升图像-文本检索准确率。与基于输出熵和余弦相似度分布的自适应策略对比,方差感知调度表现最优。t-SNE可视化显示其生成的多模态嵌入更具区分性。在注入噪声的测试中,该方法仍保持较高召回率,证明其在扰动下的鲁棒性。结果表明,自适应损失加权对低数据多模态对齐具有显著优势。

原文摘要 · Abstract (English)

Training vision-language models for image-text alignment typically requires large datasets to achieve robust performance. In low-data scenarios, standard contrastive learning can struggle to align modalities effectively due to overfitting and unstable training dynamics. In this paper, we propose a variance-aware loss scheduling approach that dynamically adjusts the weighting of the contrastive loss based on the statistical variability (uncertainty) in the model's alignment predictions. Using a subset of the Flickr8k image-caption dataset to simulate limited data conditions, we demonstrate that our approach improves image-text retrieval accuracy compared to a fixed-weight baseline. We also compare against other adaptive weighting strategies (using output entropy and cosine similarity spread) and find that variance-aware scheduling provides the best overall trade-off. Qualitatively, our method yields more distinct multimodal embeddings as shown by t-SNE visualizations. Moreover, in a stress test with noise-injected captions and images, the variance-guided loss proves more robust, maintaining higher recall when random perturbations are introduced. These results highlight the benefit of adaptive loss weighting for multimodal alignment in low-data regimes.

多模态对齐低数据损失调度对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。