首个评估大模型理解阿拉伯表格数据能力的基准,填补语言与结构化数据研究空白。
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
- 构建混合流程:大模型生成+人工审核,确保数据质量
- 大模型在简单问答中表现尚可,复杂推理仍存在明显短板
- 提出自思辨自动化评估框架,效果接近人类判别
大语言模型在自然语言处理方面取得了显著进展,但在解析结构化数据(尤其是表格)方面仍受限。尽管英语表格基准广泛可用,阿拉伯语因资源稀缺及语言特性独特而严重不足。为此,我们提出AraTable——一个全新的综合性基准,用于评估大模型在阿拉伯表格数据上的推理与理解能力。该基准涵盖直接问答、事实验证和复杂推理等多种任务,覆盖多样化的阿拉伯语表格来源。采用混合流程:先由大模型生成内容,再经人工筛选与验证,确保数据高质量。初步分析显示,大模型在直接问答等简单任务中表现尚可,但在需深层推理与事实验证的任务中仍面临重大认知挑战,表明复杂表格推理仍有巨大提升空间。此外,我们提出一种完全自动化的评估框架,利用自思辨机制,性能接近人工评判。本研究提供了公开可用的资源与评估体系,有助于推动阿拉伯语结构化数据基础模型的发展。
原文摘要 · Abstract (English)
The cognitive and reasoning abilities of large language models (LLMs) have enabled remarkable progress in natural language processing. However, their performance in interpreting structured data, especially in tabular formats, remains limited. Although benchmarks for English tabular data are widely available, Arabic is still underrepresented because of the limited availability of public resources and its unique language features. To address this gap, we present AraTable, a novel and comprehensive benchmark designed to evaluate the reasoning and understanding capabilities of LLMs when applied to Arabic tabular data. AraTable consists of various evaluation tasks, such as direct question answering, fact verification, and complex reasoning, involving a wide range of Arabic tabular sources. Our methodology follows a hybrid pipeline, where initial content is generated by LLMs and subsequently filtered and verified by human experts to ensure high dataset quality. Initial analyses using AraTable show that, while LLMs perform adequately on simpler tabular tasks such as direct question answering, they continue to face significant cognitive challenges when tasks require deeper reasoning and fact verification. This indicates that there are substantial opportunities for future work to improve performance on complex tabular reasoning tasks. We also propose a fully automated evaluation framework that uses a self-deliberation mechanism and achieves performance nearly identical to that of human judges. This research provides a valuable, publicly available resource and evaluation framework that can help accelerate the development of foundational models for processing and analysing Arabic structured data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。