发现视觉语言模型在多语言下关联能力会崩溃,跨语系表现大幅下降
Vision-Language Models are Fragile Multilingual Associators

- 构建M²BIND基准,测试跨语言输入时模型的视觉-文本绑定稳定性
- 跨语族和跨文字系统下绑定强度显著下降,内部计算移至深层且因果性减弱
- 相近语言间绑定更稳定,提示全球部署需警惕多语言性能退化
视觉语言模型需将视觉实体与文本属性关联。当输入语言变化时,这些关联是否保持稳定尚不清楚。我们提出M²BIND基准,通过变换上下文和查询的语言,在多种语言间评估绑定能力。采用外部任务指标和内部因果干预双重方式检验。结果表明,绑定并非语言不变:跨语族与跨书写系统设置中出现显著绑定崩溃,模型内部绑定计算向深层迁移,因果强度减弱。密切相关语言间的关联保留较好。总体而言,研究揭示了在多语言环境中部署的视觉语言模型,其关联质量无法保证与单语评估一致。
原文摘要 · Abstract (English)
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。