arXiv:2510.24792cs.CVcs.AI2025-10被引 3

用国际教育测评题构建多语言多模态评测集,检验模型跨语言推理能力。

PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models

  • 基于PISA测试原题构建六语种并行数据集,人工提取图文内容与题型标签。
  • 小模型(<20亿参数)在非英语题上表现显著下降,几何空间推理错误率高。
  • 适合研究多语言多模态推理的学者,尤其关注模型泛化与公平性者。

视觉语言模型在多模态推理方面取得显著进展,但现有评测集受限于高质量人工标注样本不足,多数依赖大语言模型生成的合成内容。且多数数据集仅限英文,因翻译样本的人工质量保障耗时费力。为填补这一空白,我们提出PISA-Bench,一个源自国际学生评估项目(PISA)专家设计试题的多语言评测基准,覆盖80多个国家学生能力评估体系。每条样本包含人工提取的指令、问题、选项和图像,并标注题型类别,已从英文翻译成西班牙语、德语、中文、法语和意大利语,形成完整六语言平行语料库。我们在PISA-Bench上评估了前沿视觉语言模型,发现小型模型(参数量小于200亿)得分普遍偏低,非英语语种上的性能显著下降,且在空间与几何推理任务中错误率较高。通过公开数据集与评测框架,本工作为推进多语言多模态推理研究提供基础资源。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated remarkable progress in multimodal reasoning. However, existing benchmarks remain limited in terms of high-quality, human-verified examples. Many current datasets rely on synthetically generated content by large language models (LLMs). Furthermore, most datasets are limited to English, as manual quality assurance of translated samples is time-consuming and costly. To fill this gap, we introduce PISA-Bench, a multilingual benchmark derived from English examples of the expert-created PISA tests, a unified framework for the assessment of student competencies in over eighty countries. Each example consists of human-extracted instructions, questions, answer options, and images, enriched with question type categories, and has been translated from English into five additional languages (Spanish, German, Chinese, French, and Italian), resulting in a fully parallel corpus covering six languages. We evaluate state-of-the-art vision-language models on PISA-Bench and find that especially small models (<20B parameters) fail to achieve high test scores. We further find substantial performance degradation on non-English splits as well as high error-rates when models are tasked with spatial and geometric reasoning. By releasing the dataset and evaluation framework, we provide a resource for advancing research on multilingual multimodal reasoning.

多模态评测多语言视觉语言模型PISA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。