arXiv:2605.09060cs.CL2026-05

通过对比不同语言的视觉定位表现,揭示了多语言模型失败主因在文本分支。

Language-Conditioned Visual Grounding with CLIP Multilingual

论文配图:Language-Conditioned Visual Grounding with CLIP Multilingual
图 1 · 摘自论文原文
  • 固定视觉编码器,仅更换文本分支,隔离语言差异来源
  • 低资源语言在两种模型规模下均表现显著落后(差距达0.114~0.143)
  • 空间错位是主要失效模式,而非信号丢失,适合节能部署

多语言视觉-语言模型在不同语言间存在系统性性能差距,但其成因尚不明确:跨语言差异可能源于视觉编码器、文本分支或二者交互。本文通过密集多语言CLIP探测实验,保持13种语言的视觉编码器一致,仅替换XLM-RoBERTa文本分支。在11个概念、210张图像上评估两种CLIP架构(视觉参数量相差7倍,分别为~87M和~632M),使用聚类掩码交并比(IoU)、顶百分位IoU及斯皮尔曼等级相关系数与英语参考进行对比(每语言2,310对观测)。发现:第一,低资源语言(阿拉伯语、巴斯克语、卢森堡语)在两种规模下均出现结构性劣势(基线与大模型差距均>0.114,p<10^-300),表明缺陷源自文本分支;第二,放大视觉编码器虽加剧部分语言(巴斯克语Δ=-0.056,卢森堡语Δ=-0.076)的差距,却改善阿拉伯语表现(Δ=+0.033),区分出语料覆盖不足与分词器适应性差两类问题;第三,峰值相似度在各语言间保持稳定(大模型下平均比值0.94),但聚类掩码IoU显著下降,说明空间错位是主导失败模式。该方法能耗仅为3.4–3.9 Wh/1,000查询,具备高能效优势,适合作为节能型多语言部署基础。

原文摘要 · Abstract (English)

Multilingual vision-language models exhibit systematic performance gaps across languages, but the mechanism remains ambiguous: cross-language divergence could arise from the visual encoder, the text branch, or their interaction. We resolve this ambiguity through a dense multilingual CLIP probe in which the visual encoder is held identical across thirteen typologically diverse languages and only the XLM-RoBERTa text branch varies. We evaluate two CLIP architectures spanning a 7x visual-encoder scale gap (XLM-R base + ViT-B/32, ~87M visual parameters; XLM-R large + ViT-H/14, ~632M) on 11 concepts and 210 images, and quantify cross-language agreement via cluster-mask IoU, top-percentile IoU, and Spearman rank correlation against an English reference (n=2,310 paired observations per language). Three findings emerge. First, low-resource languages (Arabic, Basque, Luxembourgish) incur a structural penalty at both backbone scales (Wilcoxon HR>LR p<10^-300; cluster-mask IoU gap +0.114 at base, +0.143 at large), isolating the deficit to the text branch. Second, scaling the encoder 7x widens the gap for structural failure cases (Basque Δ=-0.056, Luxembourgish Δ=-0.076) while improving Arabic (Δ=+0.033), separating corpus-coverage from tokeniser-fertility failures. Third, peak similarity is preserved across languages (mean ratio 0.94 at large scale) while cluster-mask IoU drops sharply, identifying spatial misalignment, not signal collapse, as the dominant failure mode. At 3.4-3.9 Wh per 1,000 queries, dense-CLIP grounding is competitive with high-throughput inference budgets, positioning it as a practical substrate for energy-aware multilingual deployment.

多语言模型视觉定位能源效率文本分支

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。