为越南语图文检索设计首个基础视觉语言模型,提升跨模态对齐效果
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
- 结合对比学习与最优传输损失,增强跨模态一致性
- 在三个越南语数据集上超越CLIP,零样本提升达11.72个百分点
- 适合低资源语言多媒体系统开发者参考应用
图文检索已成为智能多媒体系统的核心组件,但现有视觉语言模型多针对高资源语言优化,在低资源语言如越南语上表现不佳。本文提出专为越南语图文检索设计的ViCLIP-OT基础模型,融合CLIP式对比学习与相似图正则化最优传输(SIGROT)损失,提升全局跨模态一致性并缓解模态差距。在三个越南语基准(UITOpenViIC、KTVIC、Crossmodal-3600)上的实验表明,该模型在域内和零样本设置下均优于CLIP与SigLIP基线。在UIT-OpenViIC上平均Recall@K达67.34%,较CLIP提升5.75个百分点;在Crossmodal-3600零样本评估中,领先CLIP达11.72个百分点。嵌入空间分析进一步验证了更好的对齐效果与更小的模态差距。结果表明,SIGROT机制为低资源语言的跨模态检索提供了有效且可扩展的解决方案,对越南语及其他非主流语言的智能多媒体系统具有实际意义。
原文摘要 · Abstract (English)
Image-text retrieval has become a fundamental component in intelligent multimedia systems; however, most existing vision-language models are optimized for highresource languages and remain suboptimal for low-resource settings such as Vietnamese. This work introduces ViCLIP-OT, a foundation vision-language model specifically designed for Vietnamese image-text retrieval. The proposed framework integrates CLIP-style contrastive learning with a Similarity-Graph Regularized Optimal Transport (SIGROT) loss to enhance global cross-modal consistency and mitigate modality gap issues. Extensive experiments on three Vietnamese benchmarks (UITOpenViIC, KTVIC, and Crossmodal-3600) demonstrate that ViCLIP-OT consistently outperforms CLIP and SigLIP baselines in both in-domain and zero-shot settings. On UIT-OpenViIC, the model achieves an average Recall@K of 67.34%, improving upon CLIP by 5.75 percentage points. In zero-shot evaluation on Crossmodal-3600, ViCLIPOT surpasses CLIP by 11.72 percentage points. Embedding-space analysis further confirms improved alignment and reduced modality gap. The results indicate that integrating SIGROT provides an effective and scalable strategy for cross-modal retrieval in low-resource languages, offering practical implications for intelligent multimedia retrieval systems in Vietnamese and other underrepresented linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。