首个多语言图表问答基准,覆盖10种语言2.6万组问答。
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
- 通过翻译+代码复用构建多语言图表数据集
- 英文模型在非英语图表上性能显著下降
- 提供训练集,可提升各类模型的多语言理解能力
图表是通用的数据表达方式,但现有图表理解评测集高度依赖英语,难以服务全球用户。为此,我们提出PolyChartQA,首个大规模多语言图表问答基准,涵盖22,606张图表和26,151组问答对,覆盖10种不同语言。该数据集通过可扩展的生成流程构建,利用大语言模型进行翻译并复用代码,辅以严格的质量控制。我们系统评估了主流视觉语言模型(LVLMs)在多语言图表理解上的表现,发现英语与其他语言之间存在显著性能差距,尤其在低资源语言上更为明显。此外,我们还推出配套的多语言训练数据集PolyChartQA-Train,基于该数据集微调模型后,在多种规模和架构的模型上均实现显著的多语言理解性能提升。本工作为开发具备全球包容性的视觉语言模型提供了坚实基础。
原文摘要 · Abstract (English)
Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. To address this limitation, we introduce PolyChartQA, the first large-scale multilingual benchmark for chart question answering, comprising 22,606 charts and 26,151 QA pairs across 10 diverse languages. PolyChartQA is constructed through a scalable pipeline that enables efficient multilingual chart generation via data translation and code reuse, supported by LLM-based translation and rigorous quality control. We systematically evaluate multilingual chart understanding with PolyChartQA on state-of-the-art LVLMs and reveal a significant performance gap between English and other languages, particularly low-resource ones. Additionally, we introduce a companion multilingual chart question answering training set, PolyChartQA-Train, on which fine-tuning LVLMs yields substantial gains in multilingual chart understanding across diverse model sizes and architectures. Together, our benchmark provides a foundation for developing globally inclusive vision-language models capable of understanding charts across diverse linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。