构建97语言的多任务表格问答基准,推动低资源语言理解研究
M3TQA: Massively Multilingual Multitask Table Question Answering
- 基于大模型翻译构建跨语言表格问答数据集
- 涵盖2916个问答对,覆盖4类推理任务
- 特别提升低资源语言模型性能,适合多语言研究者
表格数据是现实信息系统的基石,但现有表格理解研究大多局限于英语,多语言理解仍严重不足。现有多语言表格基准存在语言地理失衡问题,部分语言过度代表,且缺乏足够规模以支撑严谨的跨语言分析。为此,我们提出大规模多语言多任务表格问答框架,包含m3TQA-Instruct——一个覆盖97种语言、跨越多个语系的大型基准,涵盖被忽视和低资源语言。通过在中英文真实表格(共50张)基础上,采用由DeepSeek和GPT-4o驱动的六步大模型翻译流程,实现高保真翻译,回译验证的中位BLEU得分达60.19。该基准包含2,916个专业标注的问答对,覆盖四个任务,用于评估精细的表格推理能力。在主流大模型上的实验揭示了跨语言泛化关键洞见:合成生成的未标注问答数据可显著提升模型表现,尤其对低资源语言效果显著。m3T-Bench确立了多语言表格理解的新标准,既提供挑战性评估平台,也给出未来研究的可扩展方法。
原文摘要 · Abstract (English)
Tabular data is a fundamental component of real-world information systems, yet most research in table understanding remains confined to English, leaving multilingual comprehension significantly underexplored. Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. To address these limitations, we introduce a comprehensive framework for massively multilingual multitask table question answering, featuring m3TQA-Instruct, a large-scale benchmark spanning 97 languages across diverse language families, including underrepresented and low-resource languages. We construct m3TQA by curating 50 real-world tables in Chinese and English, then applying a robust six-step LLM-based translation pipeline powered by DeepSeek and GPT-4o, achieving high translation fidelity with a median BLEU score of 60.19 as validated through back-translation. The benchmark includes 2,916 professionally annotated question-answering pairs across four tasks designed to evaluate nuanced table reasoning capabilities. Experiments on state-of-the-art LLMs reveal critical insights into cross-lingual generalization, demonstrating that synthetically generated, unannotated QA data can significantly boost performance, particularly for low-resource languages. M3T-Bench establishes a new standard for multilingual table understanding, providing both a challenging evaluation platform and a scalable methodology for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。