arXiv:2511.13036cs.CV2025-11AAAI被引 1

用少量参数实现多语言视觉-文本对齐,提升低资源语言性能。

uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data

  • 仅训练170万参数的投影模块,冻结原有模型
  • 在5种低资源语言上显著提升检索准确率
  • 无需配对数据,适合资源匮乏的语言场景

对比语言-图像预训练(CLIP)通过大规模英语图文对,在多种视觉任务中表现出强大泛化能力。然而,其在低资源语言中的扩展受限于高质量多语言图文数据的稀缺。现有模型在跨模态3600(XM3600)基准测试中,对捷克语、芬兰语、克罗地亚语、匈牙利语和罗马尼亚语等语言的检索性能普遍偏低。为此,我们提出一种轻量级、数据高效的多语言视觉-文本对齐框架。该方法无需图像-文本对或文本-文本对,训练时冻结预训练图像编码器和多语言文本编码器,仅训练一个170万参数的紧凑投影模块,利用英文表征作为语义锚点,通过对比损失实现对齐。这种极简训练设置使模型在监督数据有限的语言中仍能实现稳健的多语言对齐。在多个多语言检索基准上的广泛评估证实了该方法的有效性,在五个表现欠佳的语言上取得显著提升。结果表明,基于锚点的参数高效对齐策略在包容性多模态学习中具有显著优势。

原文摘要 · Abstract (English)

Contrastive Language-Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English-image pairs. However, its extension to low-resource languages remains limited due to the scarcity of high-quality multilingual image-text data. Existing multilingual vision-language models exhibit consistently low retrieval performance in underrepresented languages including Czech, Finnish, Croatian, Hungarian, and Romanian on the Crossmodal-3600 (XM3600) benchmark. To address this, we propose a lightweight and data-efficient framework for multilingual vision-language alignment. Our approach requires no image-text pairs or text-text pairs and freezes both the pretrained image encoder and multilingual text encoder during training. Only a compact 1.7M-parameter projection module is trained, using a contrastive loss over English representations as semantic anchors. This minimal training setup enables robust multilingual alignment even for languages with limited supervision. Extensive evaluation across multiple multilingual retrieval benchmarks confirms the effectiveness of our method, showing significant gains in five underrepresented languages where existing models typically underperform. These findings highlight the effectiveness of our pivot-based, parameter-efficient alignment strategy for inclusive multimodal learning.

多语言视觉-语言参数效率低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。