arXiv:2602.02774cs.CL2026-02

构建埃塞俄比亚阿姆哈拉语多文化故事问答集,揭示语言模型在地域文化理解上的差距。

AmharicStoryQA: A Multicultural Story Question Answering Benchmark in Amharic

  • 基于埃塞俄比亚不同地区阿姆哈拉语叙事构建长序列问答基准
  • 发现现有大模型在区域故事理解上存在显著差异,提升不均衡
  • 适合关注低资源语言与文化多样性评估的研究者使用

随着对大语言模型多语言、跨文化评估基准的重视,语言与文化常被等同看待,性能常被视为模型理解语言的代理指标。本文认为此类评估忽略了同一语言内部存在的显著文化差异。为此,我们聚焦埃塞俄比亚不同地区的叙事内容,证明尽管语言特征共享,但地域与领域特异性内容会显著影响语言评估结果。我们引入 extbf{ extit{AmharicStoryQA}},一个基于阿姆哈拉语多文化叙事的长序列故事问答基准。通过该基准,我们揭示现有大模型存在显著的叙事理解差距,发现评估结果呈现明显区域差异,且监督微调在不同区域和场景中的提升效果不均。研究强调需要超越语言层面、基于文化背景的评估基准,以更准确评估并改进低资源语言的叙事理解能力。

原文摘要 · Abstract (English)

With the growing emphasis on multilingual and cultural evaluation benchmarks for large language models, language and culture are often treated as synonymous, and performance is commonly used as a proxy for a models understanding of a given language. In this work, we argue that such evaluations overlook meaningful cultural variation that exists within a single language. We address this gap by focusing on narratives from different regions of Ethiopia and demonstrate that, despite shared linguistic characteristics, region-specific and domain-specific content substantially influences language evaluation outcomes. To this end, we introduce \textbf{\textit{AmharicStoryQA}}, a long-sequence story question answering benchmark grounded in culturally diverse narratives from Amharic-speaking regions. Using this benchmark, we reveal a significant narrative understanding gap in existing LLMs, highlight pronounced regional differences in evaluation results, and show that supervised fine-tuning yields uneven improvements across regions and evaluation settings. Our findings emphasize the need for culturally grounded benchmarks that go beyond language-level evaluation to more accurately assess and improve narrative understanding in low-resource languages.

多文化评估阿姆哈拉语叙事理解低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。