arXiv:2508.03737cs.CLcs.AI2025-08中稿 · , Presented and Pu…被引 1

构建中英双语数学推理基准,评估视觉语言模型在真实考试题上的表现。

GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models

  • 基于印度两大考试真题,设计含图数学题的双语测试集。
  • GPT-4o mini在零样本和两样本链式思维下最高准确率38.15%。
  • 模型在印地语答题时性能下降,双锁约束更凸显推理短板。

近年来,针对视觉语言模型(VLMs)在多个领域推理能力的评测基准日益增多,但多为英文单语。此外,除理解与翻译外,缺乏其他任务的印地语数据集。本文提出GanitBench,一个包含1527个仅图像问题的双语数学推理基准,覆盖多个数学主题,提供英语与印地语版本。数据源自印度两大重要考试——JEE高级考试与CBSE高中考试,题目以图像形式呈现,包含解题必需的图表与文字。我们在零样本链式思维(CoT)与两样本CoT设置下评估两个闭源模型。结果显示,GPT-4o mini表现更优,最高平均准确率为38.15%。通过“双锁”约束(Double Lock)测试,模型性能显著下降,两样本CoT在此环境下更有效。模型在印地语任务中的表现也低于英语,表明语言差异影响推理能力。本工作旨在推动印地语等非英语语言在视觉语言模型研究中的应用。

原文摘要 · Abstract (English)

Benchmarks for evaluating reasoning among Vision Language Models (VLMs) on several fields and domains are being curated more frequently over the last few years. However these are often monolingual, mostly available in English. Additionally there also is a lack of datasets available in Hindi on tasks apart from comprehension and translation. We introduce GanitBench, a tough benchmark consisting of 1527 vision-only questions covering several topics in Mathematics - available in languages English and Hindi. Collected from two major examinations from India, the JEE Advanced and the CBSE Boards examinations, this benchmark includes questions in the form of images comprising of figures essential to a question as well as text. We evaluate two closed source models for the same, in zero-shot Chain-of-Thought (CoT) and two-shot CoT settings. GPT-4o mini is found to be the more dominant model on the benchmark, with it's highest average accuracy being 38.15%. We also evaluate models through a "Double Lock" constraint, which brings down the performance of the models by considerable margins. We observe that two-shot CoT appears to be a more effective setting under this environment. Performance of the two VLMs also decreases when answering the same questions in the Hindi language. We hope to facilitate the inclusion of languages like Hindi in research through our work.

数学推理双语评测视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。