arXiv:2410.03576cs.CL2024-10EMNLP被引 13

为低资源印地语和孟加拉语构建大规模自动表问答数据集

Table Question Answering for Low-resourced Indic Languages

  • 基于自动数据生成技术,无需人工标注
  • 模型在新数据集上超越现有大模型表现
  • 适合对低资源语言智能系统研究者参考

表问答(TableQA)是针对结构化表格回答问题的任务,输出为单个单元格或表格。现有研究主要集中在高资源语言,导致中低资源语言进展缓慢,原因在于标注数据与神经模型的匮乏。本文提出一种全自动的大规模表问答数据生成方法,应用于两种无表问答数据集和模型的印地语与孟加拉语。基于该方法构建的数据集训练的模型性能超越当前最先进的大语言模型。我们进一步评估了模型在数学推理能力和零样本跨语言迁移方面的表现。本工作首次聚焦于可扩展的数据生成与评估流程,适用于任何具有网络存在的低资源语言。相关数据集、模型与代码已开源(https://github.com/kolk/Low-Resource-TableQA-Indic-languages)。

原文摘要 · Abstract (English)

TableQA is the task of answering questions over tables of structured information, returning individual cells or tables as output. TableQA research has focused primarily on high-resource languages, leaving medium- and low-resource languages with little progress due to scarcity of annotated data and neural models. We address this gap by introducing a fully automatic large-scale tableQA data generation process for low-resource languages with limited budget. We incorporate our data generation method on two Indic languages, Bengali and Hindi, which have no tableQA datasets or models. TableQA models trained on our large-scale datasets outperform state-of-the-art LLMs. We further study the trained models on different aspects, including mathematical reasoning capabilities and zero-shot cross-lingual transfer. Our work is the first on low-resource tableQA focusing on scalable data generation and evaluation procedures. Our proposed data generation method can be applied to any low-resource language with a web presence. We release datasets, models, and code (https://github.com/kolk/Low-Resource-TableQA-Indic-languages).

表问答低资源语言数据生成印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。