arXiv:2510.21606cs.CV2025-10

用轻量方法提升视觉语言模型在少样本下的对齐能力

Modest-Align: Data-Efficient Alignment for Vision-Language Models

  • 引入随机扰动和嵌入平滑缓解弱对齐数据的过自信
  • 在100倍少数据、600倍少显存下超越CLIP性能
  • 适合资源受限场景下的跨模态对齐应用

跨模态对齐旨在将异构模态映射到共享潜在空间,如CLIP模型通过大规模图文预训练获得强大识别能力。但在资源受限环境下,因数据量少或质量差,存在大量模糊或弱相关图文对,导致模型过自信且性能下降。现有对比学习方法依赖单一正样本对,进一步强化了对不确定样本的过自信。为此,我们提出Modest-Align,一种轻量级对齐框架,具备鲁棒性与高效性。该方法结合两种互补策略:随机扰动(引入可控噪声模拟不确定性)与嵌入平滑(校准嵌入空间中的相似度分布),共同降低过自信并提升在噪声或弱对齐样本上的表现。在多个基准数据集上的大量实验表明,Modest-Align在检索任务中优于现有最先进方法,在超过100倍更少训练数据和600倍更少GPU时间条件下仍保持竞争力。本方法为真实世界低资源场景下的跨模态对齐提供了实用且可扩展的解决方案。

原文摘要 · Abstract (English)

Cross-modal alignment aims to map heterogeneous modalities into a shared latent space, as exemplified by models like CLIP, which benefit from large-scale image-text pretraining for strong recognition capabilities. However, when operating in resource-constrained settings with limited or low-quality data, these models often suffer from overconfidence and degraded performance due to the prevalence of ambiguous or weakly correlated image-text pairs. Current contrastive learning approaches, which rely on single positive pairs, further exacerbate this issue by reinforcing overconfidence on uncertain samples. To address these challenges, we propose Modest-Align, a lightweight alignment framework designed for robustness and efficiency. Our approach leverages two complementary strategies -- Random Perturbation, which introduces controlled noise to simulate uncertainty, and Embedding Smoothing, which calibrates similarity distributions in the embedding space. These mechanisms collectively reduce overconfidence and improve performance on noisy or weakly aligned samples. Extensive experiments across multiple benchmark datasets demonstrate that Modest-Align outperforms state-of-the-art methods in retrieval tasks, achieving competitive results with over 100x less training data and 600x less GPU time than CLIP. Our method offers a practical and scalable solution for cross-modal alignment in real-world, low-resource scenarios.

跨模态对齐少样本学习模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。