为印地语和泰卢固语构建首个多语言对齐基准,揭示主流模型在此类语言上性能下降。
HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples
- 结合回译、过滤与人工验证,半自动化构建跨语言数据集。
- 每种语言约4000个问答对,覆盖科学、文化等本土场景。
- 发现5大模型在印地语/泰卢固语上平均性能下降8.3/5.5分,凸显多模态理解短板。
印度拥有近15亿人口和超过120种主要语言,是全球最多样化的地区之一。随着多语言视觉-语言模型(VLMs)日益重要,针对低资源语言的可靠评估方法亟需建立。现有评估存在四大缺陷:依赖未经验证的自动翻译、任务与领域覆盖窄、样本量有限、缺乏文化和原生来源的问答数据。为此,我们提出一个可扩展的框架,用于评估印度语言中的VLM表现,并与英语表现对比。基于该框架,我们构建了HinTel-AlignBench基准,数据源自印地语和泰卢固语的多样化来源,并配有英文对齐样本。贡献包括:(1) 一种融合回译、过滤与人工验证的半自动化数据构建框架;(2) 当前最全面的印地语与泰卢固语视觉-语言基准,包含适配的英文数据集(VQAv2, RealWorldQA, CLEVR-Math)及本土原创数据集(JEE用于理工科,VAANI用于文化理解),每种语言约4000个问答对;(3) 对多种前沿开源与闭源VLM的详细性能分析。结果显示,5个模型中4个在印度语言任务上的表现均低于英语,印地语平均下降8.3分,泰卢固语下降5.5分。我们归纳出常见失败模式,指明多语言多模态理解的关键改进方向。
原文摘要 · Abstract (English)
With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to drive progress toward equitable AI for low-resource languages. Current multilingual VLM evaluations suffer from four major limitations: reliance on unverified auto-translations, narrow task/domain coverage, limited sample sizes, and lack of cultural and natively sourced Question-Answering (QA). To address these gaps, we present a scalable framework to evaluate VLMs in Indian languages and compare it with performance in English. Using the framework, we generate HinTel-AlignBench, a benchmark that draws from diverse sources in Hindi and Telugu with English-aligned samples. Our contributions are threefold: (1) a semi-automated dataset creation framework combining back-translation, filtering, and human verification; (2) the most comprehensive vision-language benchmark for Hindi and and Telugu, including adapted English datasets (VQAv2, RealWorldQA, CLEVR-Math) and native novel Indic datasets (JEE for STEM, VAANI for cultural grounding) with approximately 4,000 QA pairs per language; and (3) a detailed performance analysis of various State-of-the-Art (SOTA) open-weight and closed-source VLMs. We find a regression in performance for tasks in English versus in Indian languages for 4 out of 5 tasks across all the models, with an average regression of 8.3 points in Hindi and 5.5 points for Telugu. We categorize common failure modes to highlight concrete areas of improvement in multilingual multimodal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。