测试多文字语言模型发现,同一语言不同书写系统表现差异大,模型不真正支持多文字。
Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

- 构建包含三种旁遮普文字的平行图文数据集,检验模型跨文字一致性。
- 模型在不同文字间准确率差距达16%,视觉信息无法弥合文字鸿沟。
- 提出新评估指标SCR,最低仅24.8%,呼吁公平多文字评估标准。
当前视觉语言模型(VLMs)的多语言评估假设语言与书写系统一一对应,忽视了数十亿使用多文字语言的用户。我们提出PuMVR(旁遮普多模态视觉推理)基准,包含1,000个严格平行的图像-文本实例,覆盖旁遮普语的三种活跃书写系统:古木基文、沙姆基文和罗马字母。评估10个顶尖VLMs发现存在显著且系统性的「文字差距」:模型在一种文字中能完成视觉任务,但在另一文字中却失败,准确率差距最高达16%。关键的是,视觉输入虽能统一提升绝对性能,但无法缩小书写系统间的差距。此外,跨文字上下文迁移极为脆弱,暴露模型存在文字锁定的知识表示。通过所有文字对的McNemar检验支持,结果表明当前所谓的“多语言”VLMs并非真正多文字。我们提出文字一致性率(SCR),在本基准上最低为24.8%,建议其作为实现无文字偏见评估的强制指标。数据与代码已公开于:https://github.com/prabhjotschugh/Not-Truly-Multilingual-PuMVR。
原文摘要 · Abstract (English)
Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, we expose a substantial and systematic Script Gap. Models frequently solve visual tasks in one script while failing identical tasks in another, with accuracy deltas reaching 16%. Crucially, visual input boosts absolute performance uniformly yet does not close the orthographic gap. Furthermore, cross-script in-context transfer is highly brittle, exposing script-locked knowledge representation. Supported by McNemar tests across all script pairs, our findings demonstrate that current "multilingual" VLMs are not truly multi-script. We propose the Script Consistency Rate (SCR), which falls as low as 24.8% on our benchmark, as a mandatory metric for script-agnostic evaluation to ensure equitable AI access. Data and code are available at: https://github.com/prabhjotschugh/Not-Truly-Multilingual-PuMVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。