构建首个日文图表表格理解基准,推动多语言视觉语言模型发展
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

- 从33份政府白皮书提取真实图表,构建大规模日文图文问答数据集
- 含2053张图像、10种以上图类型,人工标注问答对评估深层理解能力
- 开源数据集与代码,揭示现有模型在复杂图表理解上仍有巨大提升空间
理解图表和表格图像对于将视觉语言模型(VLMs)应用于真实文档理解至关重要。尽管英语基准发展迅速,但非英语基准仍极为稀缺,难以判断该进展是否具备跨语言泛化能力。主要障碍在于大规模收集真实且多样化的非英语图表和表格图像的难度。为此,我们利用政府白皮书作为可扩展的数据源,超越英语语境,因其包含跨领域、多格式的自然出现图表和表格,且在许多国家免费可获取。作为首个实例,我们提出HakushoBench,一个基于33份政府白皮书构建的具有挑战性的日文图表与表格视觉问答基准。该数据集包含2053张图像,涵盖10余种图像类型,配有手动标注的问答对,旨在评估对图表和表格的深度与整体理解,而非仅依赖局部视觉线索。在多种VLM上的实验表明,该基准对开源模型依然极具挑战:表现最佳的开源模型准确率仅为58.6%,而开源与专有模型之间存在34.9个百分点的差距,凸显了在复杂图表理解方面仍有巨大改进空间。我们已公开数据集与代码。
原文摘要 · Abstract (English)
Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchmarks have advanced rapidly, non-English counterparts remain scarce, leaving it unclear whether this progress generalizes across languages. A key obstacle is the difficulty of collecting realistic and diverse non-English chart and table images at scale. To address this, we leverage governmental white papers as a scalable source for benchmark construction beyond English, as they contain naturally occurring charts and tables across diverse formats and domains and are freely accessible in many countries. As a first instantiation, we introduce HakushoBench, a challenging Japanese chart and table VQA benchmark built from 33 governmental white papers. HakushoBench contains 2,053 images spanning over 10 image types, with manually annotated QA pairs, designed to assess deep and holistic understanding of charts and tables, rather than local visual cues alone. Experiments across a broad range of VLMs demonstrate that HakushoBench remains challenging for open-weight models: the best open-weight model achieves only 58.6% accuracy, and a 34.9-point gap between open-weight and proprietary models highlights substantial room for improvement in complex chart and table understanding. We release our dataset and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。