arXiv:2503.00231cs.CLcs.AI2025-03被引 24

构建阿拉伯谚语多方言数据集,评估大模型跨文化理解能力

Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

  • 创建涵盖多种阿拉伯方言的谚语数据集Jawaher,含意译与解释
  • 大模型能生成准确译文,但难给出文化语境下的合理解释
  • 适合研究多语言、跨文化语言理解的学者与开发者

近期指令微调、基于人类反馈的强化学习(RLHF)及直接偏好优化(DPO)等技术显著提升了大语言模型(LLMs)对用户偏好的适应性。然而,许多模型仍存在对西方、英语中心或美国文化的偏见,其在英语数据上的表现持续优于其他语言,暴露出模型在文化多样性方面的系统性差距,尤其在处理如谚语这类富含文化内涵的隐喻语言时表现不足。为此,我们提出Jawaher——一个用于评估大模型理解与解读阿拉伯谚语能力的基准数据集。该数据集包含来自不同阿拉伯方言的谚语,附带惯用译法和解释说明。通过对开源与闭源模型的广泛评测发现,尽管模型能生成语义准确的翻译,但在生成文化背景精准且上下文相关的解释方面仍存在明显短板。这表明需持续优化模型并扩展数据集,以弥合大模型在隐喻语言处理中的文化鸿沟。

原文摘要 · Abstract (English)

Recent advancements in instruction fine-tuning, alignment methods such as reinforcement learning from human feedback (RLHF), and optimization techniques like direct preference optimization (DPO) have significantly enhanced the adaptability of large language models (LLMs) to user preferences. However, despite these innovations, many LLMs continue to exhibit biases toward Western, Anglo-centric, or American cultures, with performance on English data consistently surpassing that of other languages. This reveals a persistent cultural gap in LLMs, which complicates their ability to accurately process culturally rich and diverse figurative language such as proverbs. To address this, we introduce Jawaher, a benchmark designed to assess LLMs' capacity to comprehend and interpret Arabic proverbs. Jawaher includes proverbs from various Arabic dialects, along with idiomatic translations and explanations. Through extensive evaluations of both open- and closed-source models, we find that while LLMs can generate idiomatically accurate translations, they struggle with producing culturally nuanced and contextually relevant explanations. These findings highlight the need for ongoing model refinement and dataset expansion to bridge the cultural gap in figurative language processing.

语言模型文化差异谚语理解多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。