arXiv:2606.20770cs.CLcs.AI2026-06中稿 · the 4th Workshop o…

首个量化多文字语言模型偏见的基准,发现同一语言不同书写系统下模型表现差异高达16%。

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

  • 设计跨三种文字的图像推理任务,测试模型在旁遮普语中的表现差异。
  • 10个先进模型在不同文字间准确率差距达16%,一致性最低仅24.8%。
  • 提出新评估指标SCR,推动公平多语言AI发展,适合多语言研究者参考。

当前视觉语言模型(VLMs)虽具备多语言能力,但默认一语言对应单一书写系统,忽视了如旁遮普语、塞尔维亚语、印地-乌尔都语等多文字语言的数亿用户。我们提出PuMVR(旁遮普语多模态视觉推理)基准,通过375个文化相关的图像推理任务,评估旁遮普语三种活跃书写系统(Gurmukhi、Shahmukhi、Roman)下的模型表现。测试10个前沿VLMs发现显著的‘文字差距’:模型在一种文字中能解决视觉谜题,却在另一文字中失败,准确率差异最高达16%,文字一致性率(SCR)低至24.8%。视觉输入虽提升整体性能,但无法消除相对偏差。分析显示推理路径受文字影响,跨文字迁移能力弱。我们提出以SCR为核心指标,挑战现有多语言评估范式,为实现无文字偏见的公平人工智能提供框架。

原文摘要 · Abstract (English)

Current Vision-Language Models (VLMs) are celebrated for their multilingual capabilities, yet they operate under a flawed assumption: that one language corresponds to a single writing system. This overlooks billions of users of multi-script languages like Punjabi, Serbian, Hindi-Urdu, Kurdish, among many others, for whom a model's capability may be fractured by orthographic bias. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), the first benchmark designed to quantify script-dependent bias through 375 culturally grounded image-reasoning tasks across Punjabi's three active scripts (Gurmukhi, Shahmukhi, Roman). Evaluating 10 state-of-the-art VLMs, we expose a substantial Script Gap: models frequently solve visual puzzles in one script while failing identical tasks in another, with accuracy deltas reaching 16% and Script Consistency Rates (SCR) as low as 24.8%. Crucially, visual input boosts absolute performance but does not close this gap, the relative bias persists. Our analysis suggests reasoning patterns show limited cross-script transferability, and Chain-of-Thought pathways diverge based on script alone. We propose SCR as a core metric for script-agnostic evaluation, challenging current multilingual assessment paradigms and providing a framework for equitable AI.

多语言视觉语言模型文字偏见公平AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。