首次评测多语言视觉模型在印度语种上的推理能力,发现中文外语言表现显著下降。
Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
- 将三类英文推理题翻译成六种印度语言,用双模型交叉验证确保质量。
- 模型在印度语言上准确率比英语低9.8至25个百分点,德拉维达语种降幅更大。
- 英文推理链反噬本地语言,提示需为非英语设计专用推理机制。
视觉语言模型在数学、科学和空间推理基准上表现优异,但这些评估几乎全为英语。本文首次开展针对印度语言的跨语言视觉推理审计。将来自 MathVista、ScienceQA 和 MMMU 的 980 道题目通过 IndicTrans2 翻译为印地语、泰米尔语、泰卢固语、孟加拉语、卡纳达语和马拉地语,并用 Gemini 2.0 Flash 对每种语言的 50 个样本进行交叉验证(译者间一致性 0.79–0.84)。评估了八种视觉语言模型(从 7B 开源模型到 GPT-4o),覆盖全部七种语言,生成 68,600 条推理记录,包含纯文本与思维链消融实验。结果发现,从英语切换至印度语言后,准确率下降 9.8–25 个百分点,其中德拉维达语种比印欧语种额外下降最多达 13.2 个百分点。思维链提示对孟加拉语(-14.4 个百分点)和卡纳达语(-11.4 个百分点)产生负面影响,暴露其英文中心的推理路径缺陷。即使专为 23 种语言设计的 Aya-Vision-8B 模型,在德拉维达文字上仍下降 28.5 个百分点,表明多语言预训练不足以迁移视觉推理能力。研究释放了翻译后的基准及所有模型输出。
原文摘要 · Abstract (English)
Vision-language models score well on mathematical, scientific, and spatial reasoning benchmarks, yet these evaluations are overwhelmingly English. I present the first cross-lingual visual reasoning audit for Indian languages. 980 questions from MathVista, ScienceQA, and MMMU are translated into Hindi, Tamil, Telugu, Bengali, Kannada, and Marathi using IndicTrans2, with Gemini 2.0 Flash cross-verification on 50 samples per language (inter-translator agreement 0.79-0.84). Eight VLMs, from 7B open-source models to GPT-4o, are evaluated across all seven languages, yielding 68,600 inference records that include text-only and chain-of-thought ablations. I find accuracy drops of 9.8-25 percentage points when switching from English to an Indian language, with Dravidian languages suffering up to 13.2 pp more than Indo-Aryan. Chain-of-thought prompting degrades Bengali (-14.4 pp) and Kannada (-11.4 pp) rather than helping, exposing English-centric reasoning chains. Aya-Vision-8B, built for 23 languages, still drops 28.5 pp on Dravidian scripts; multilingual pretraining alone does not transfer visual reasoning. I release the translated benchmark and all model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。