arXiv:2508.16139cs.CL2025-08EMNLP被引 3

构建首个关注文化地域差异的多语言开放域问答基准

XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering

  • 基于3000个英文问题跨8语言扩展,确保语义一致
  • 发现主流模型在地域敏感问题上表现显著下降
  • 适合评估多语言系统在真实文化场景下的能力

大型语言模型在开放域问答(ODQA)上取得显著进展,但现有评估主要聚焦英文,且假设答案在不同语言间无地域差异。这一假设忽略了文化与地区差异对问题理解与答案的影响,导致多语言基准存在偏见。为此,我们提出XLQA,一个专为地域敏感型多语言ODQA设计的新基准。该数据集包含3000个英文种子问题,扩展至八种语言,并通过严格筛选保证语义一致性,人工验证标注了地域不变与地域敏感两类问题。对五种先进多语言大模型的评估显示,其在地域敏感问题上表现明显不足,暴露出英语与其他语言之间因缺乏地域知识而导致的能力差距。我们提供系统性框架与可扩展方法,用于在多元文化背景下评估多语言问答,为提升多语言问答系统的真实应用能力提供关键资源。研究发现,训练数据分布差异是导致模型语言能力与地域感知能力差异的重要原因。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown significant progress in Open-domain question answering (ODQA), yet most evaluations focus on English and assume locale-invariant answers across languages. This assumption neglects the cultural and regional variations that affect question understanding and answer, leading to biased evaluation in multilingual benchmarks. To address these limitations, we introduce XLQA, a novel benchmark explicitly designed for locale-sensitive multilingual ODQA. XLQA contains 3,000 English seed questions expanded to eight languages, with careful filtering for semantic consistency and human-verified annotations distinguishing locale-invariant and locale-sensitive cases. Our evaluation of five state-of-the-art multilingual LLMs reveals notable failures on locale-sensitive questions, exposing gaps between English and other languages due to a lack of locale-grounding knowledge. We provide a systematic framework and scalable methodology for assessing multilingual QA under diverse cultural contexts, offering a critical resource to advance the real-world applicability of multilingual ODQA systems. Our findings suggest that disparities in training data distribution contribute to differences in both linguistic competence and locale-awareness across models.

多语言问答地域敏感评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。