arXiv:2412.04193cs.CL2024-12被引 18

系统评估大模型在阿拉伯方言中的表现,发现其理解强但不愿生成方言。

AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic

  • 构建四维度评估框架,覆盖方言保真度、理解力、生成质量与双语差异。
  • 9个大模型在8种阿拉伯方言中测试,显示生成能力显著弱于理解能力。
  • 提示词优化可缓解方言生成偏见,适合关注公平语言技术的研究者。

阿拉伯方言(DA)在语言技术中长期被忽视,尤其是大语言模型(LLMs)的应用严重不足,这可能加剧社会不平等并限制模型实际应用。然而,研究界缺乏对阿拉伯方言性能的可操作评估标准。本文提出一个涵盖保真度、理解力、质量与双语差异四个维度的综合评估框架,对九个主流大模型在八种阿拉伯方言上的表现进行了系统评估,并提供实践建议。结果表明,尽管大模型能较好理解阿拉伯方言,但在生成方面表现较弱,非因方言能力差,而是出于生成意愿低。进一步分析发现,当前的后训练过程可能加剧对阿拉伯方言的偏见;少量示例可有效缓解此问题;而输入文本特征与模型方言表现之间无显著相关性。

原文摘要 · Abstract (English)

Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs). This trend threatens to exacerbate existing social inequalities and limits LLM applications, yet the research community lacks operationalized performance measurements in DA. We present a framework that comprehensively assesses LLMs' DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia. We evaluate nine LLMs in eight DA varieties and provide practical recommendations. Our evaluation suggests that LLMs do not produce DA as well as they understand it, not because their DA fluency is poor, but because they are reluctant to generate DA. Further analysis suggests that current post-training can contribute to bias against DA, that few-shot examples can overcome this deficiency, and that otherwise no measurable features of input text correlate well with LLM DA performance.

大模型评估阿拉伯语方言生成公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。