arXiv:2508.17334cs.CVcs.AI2025-08被引 3

测试大模型在跨语言板球数据上的数字与语言推理能力

Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

  • 构建板球记分表图文问答数据集,含英印双语格式
  • 顶尖模型在英文任务上表现不佳,跨语言时更差
  • 适合研究视觉语言模型结构理解与跨语言泛化能力

我们提出MMCRICBENCH-3K,一个针对板球记分表的视觉问答基准,用于评估大视觉语言模型(LVLMs)在半结构化表格图像上的复杂数值与跨语言推理能力。该数据集包含1,463张由ODI、T20和测试赛格式合成的记分表图像,配套1,500个英文问答对,分为两个子集:MMCRICBENCH-E-1.5K(英文记分表)和MMCRICBENCH-H-1.5K(视觉相似的印地语记分表),所有问题与答案均保留英文,以实现受控的跨脚本评估。任务要求对结构化数值数据、多图上下文及隐含领域知识进行推理。实证结果显示,即使最先进的模型如GPT-4o和Qwen2.5VL,在其主要训练语言的英文子集上仍表现不佳,且在印地语子集上性能进一步下降。这揭示了现有模型在结构感知视觉文本理解、数值推理和跨语言泛化方面的关键局限。数据集已通过Hugging Face公开:https://huggingface.co/datasets/DIALab/MMCricBench,以推动该方向的研究。

原文摘要 · Abstract (English)

We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-structured tabular images. MMCRICBENCH-3K comprises 1,463 synthetically generated scorecard images from ODI, T20, and Test formats, accompanied by 1,500 English QA pairs. It includes two subsets: MMCRICBENCH-E-1.5K, featuring English scorecards, and MMCRICBENCH-H-1.5K, containing visually similar Hindi scorecards, with all questions and answers kept in English to enable controlled cross-script evaluation. The task demands reasoning over structured numerical data, multi-image context, and implicit domain knowledge. Empirical results show that even state-of-the-art LVLMs, such as GPT-4o and Qwen2.5VL, struggle on the English subset despite it being their primary training language and exhibit a further drop in performance on the Hindi subset. This reveals key limitations in structure-aware visual text understanding, numerical reasoning, and cross-lingual generalization. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/DIALab/MMCricBench, to promote LVLM research in this direction.

视觉语言模型跨语言数值推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。