arXiv:2508.13130cs.CL2025-08

首个阿拉伯多方言常识验证数据集及推理方法

MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation

  • 构建多方言阿拉伯语常识数据集,覆盖口语化表达
  • 用图卷积网络提升语义关系建模,准确率显著优于基线
  • 适合研究阿拉伯语自然语言理解与跨方言模型的学者

常识验证旨在判断句子是否符合日常人类认知,是构建稳健自然语言理解系统的关键能力。尽管英语领域已取得显著进展,阿拉伯语相关研究仍不充分,尤其受限于其丰富的语言多样性。现有资源多集中于现代标准阿拉伯语(MSA),而方言在口语中更为普遍,却长期被忽视。为此,我们提出两项核心贡献:首次构建涵盖多种方言的阿拉伯语常识推理数据集MuDRiC;并提出一种新方法,将图卷积网络(GCNs)适配至阿拉伯语常识推理任务,增强语义关系建模能力。实验表明,该方法持续优于直接微调语言模型的基线。本工作为阿拉伯语自然语言理解提供了基础数据集与新方法,有效应对语言变体复杂性。数据与代码见https://github.com/KareemElozeiri/MuDRiC。

原文摘要 · Abstract (English)

Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. While substantial progress has been made in English, the task remains underexplored in Arabic, particularly given its rich linguistic diversity. Existing Arabic resources have primarily focused on Modern Standard Arabic (MSA), leaving regional dialects underrepresented despite their prevalence in spoken contexts. To bridge this gap, we present two key contributions. We introduce MuDRiC, an extended Arabic commonsense dataset incorporating multiple dialects. To the best of our knowledge, this is the first Arabic multi-dialect commonsense reasoning dataset. We further propose a novel method adapting Graph Convolutional Networks (GCNs) to Arabic commonsense reasoning, which enhances semantic relationship modeling for improved commonsense validation. Our experimental results demonstrate that this approach consistently outperforms the baseline of direct language model fine-tuning. Overall, our work enhances Arabic natural language understanding by providing a foundational dataset and a new method for handling its complex variations. Data and code are available at https://github.com/KareemElozeiri/MuDRiC.

常识推理阿拉伯语多方言图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。