首个针对低资源语言的多模态理解评测,揭示模型在非英语场景下的能力短板。
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
- 构建八任务多语言基准VLURes,涵盖长文本与跨语言视觉语言理解。
- GPT-4o在多语言上达90.8%准确率,但开源模型差距显著,人类仍领先6.7%。
- 首次为斯瓦希里语和乌尔都语提供高质量图文数据集,推动公平评估。
视觉语言模型(VLMs)对智能体感知能力的发展至关重要,但现有评估大多局限于以英语为主的短文本图像-文本配对基准。为评估VLM在长文本设定下四种语言的细粒度理解能力,我们提出新型多语言基准VLURes,包含八项视觉语言任务及首创的无关性任务,用于探测英语、日语以及低资源语言斯瓦希里语和乌尔都语中VLM的细粒度视觉与语言理解能力。数据集从目标语言网络资源中构建,涵盖十类图像类别和丰富文本上下文,为斯瓦希里语和乌尔都语提供了宝贵的视觉语言资源。通过提示VLM生成回答与推理过程,并结合自动评估与母语者评分,我们发现语言与任务间存在显著性能差异,影响物体识别、场景理解与关系理解等关键智能体能力。我们对十种VLM进行了评估,表现最佳的GPT-4o在总体上达到90.8%准确率,仅比人类低6.7%,而开源模型差距更大。该结果凸显了VLURes在推动多模态视觉推理智能体发展中的关键作用。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in which the image-text pairs comprise short texts. To evaluate VLM fine-grained abilities, in four languages under long-text settings, we introduce a novel multilingual benchmark VLURes featuring eight vision-and-language tasks, and a pioneering unrelatedness task, to probe the fine-grained Visual and Linguistic Understanding capabilities of VLMs across English, Japanese, and low-resource languages, Swahili, and Urdu. Our datasets, curated from web resources in the target language, encompass ten diverse image categories and rich textual context, introducing valuable vision-language resources for Swahili and Urdu. By prompting VLMs to generate responses and rationales, evaluated automatically and by native speakers, we uncover performance disparities across languages and tasks critical to intelligent agents, such as object recognition, scene understanding, and relationship understanding. We conducted evaluations of ten VLMs with VLURes. The best performing model, GPT-4o, achieves an overall accuracy of 90.8% and lags human performance by 6.7%, though the gap is larger for open-source models. The gap highlights VLURes' critical role in developing intelligent agents to tackle multi-modal visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。