arXiv:2502.17253cs.CL2025-02EMNLP被引 2

首个多语言表格文本问答数据集,揭示模型在非英语上的性能下降19.4%。

MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering

  • 从3个英文数据集采样并翻译成10种语言构建MULTITAT
  • 非英语数据上模型平均性能下降19.4%,凸显多语言挑战
  • 提出新基线模型,平均优于其他方法3.3分,适合多语言研究者

表格与文本联合问答(TATQA)是数据密集型领域的重要任务,但现有数据集仅限于英文,存在两大缺陷:一是忽略多语言场景下的挑战,无法评估模型在多语言环境中的表现;二是未能反映真实世界中表格和文本常以非英语出现的现实。为此,我们提出了首个多语言TATQA数据集MULTITAT,从3个主流TATQA数据集中采样并翻译为10种不同语言。为对齐模型在英文与其他语言上的问答能力,我们构建了基线模型Ours。实验表明,非英语数据上性能平均下降19.4%,证明了构建MULTITAT的必要性。进一步分析揭示性能差距原因。Ours相比其他基线平均提升3.3分,验证其有效性。

原文摘要 · Abstract (English)

Question answering on the hybrid context of tables and text (TATQA) is a critical task, with broad applications in data-intensive domains. However, existing TATQA datasets are limited to English, leading to several drawbacks: (i) They overlook the challenges of multilingual TAT-QA and cannot assess model performance in the multilingual setting. (ii) They do not reflect real-world scenarios where tables and texts frequently appear in non-English languages. To address the limitations, we propose the first multilingual TATQA dataset (MULTITAT). Specifically, we sample data from 3 mainstream TATQA datasets and translate it into 10 diverse languages. To align the model TATQA capabilities in English with other languages, we develop a baseline, Ours. Experimental results reveal that the performance on non-English data in MULTITAT drops by an average of 19.4% compared to English, proving the necessity of MULTITAT. We further analyze the reasons for this performance gap. Furthermore, Ours outperforms other baselines by an average of 3.3, demonstrating its effectiveness.

表格问答多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。