arXiv:2507.04014cs.CLcs.AI2025-07ACL被引 5

评测大模型对韩国民俗迷信的文化推理能力,发现语言提示影响巨大。

Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition

  • 构建247题韩国民俗基准测试,覆盖31个主题。
  • 模型虽识事实但难实用,语言提示显著提升表现。
  • 适合研究多文化语境下AI认知偏差的学者使用。

随着大语言模型在各领域成为关键顾问,其文化敏感性和推理能力在多元文化环境中至关重要。我们提出Nunchi-Bench,一个聚焦韩国民俗迷信的文化理解评测基准,包含247道题,覆盖31个主题,评估事实知识、文化适宜建议及情境解读能力。我们在韩语和英语双语环境下评估多语言大模型,分析其对韩国文化语境的推理表现及语言差异的影响。为此,我们设计了定制化评分策略,量化模型对文化细微差别的识别与恰当回应程度。结果表明,模型虽能识别事实信息,但在实际场景中应用能力不足;明确的文化提示比仅依赖语言提示更有效。为支持后续研究,我们公开发布Nunchi-Bench及排行榜。

原文摘要 · Abstract (English)

As large language models (LLMs) become key advisors in various domains, their cultural sensitivity and reasoning skills are crucial in multicultural environments. We introduce Nunchi-Bench, a benchmark designed to evaluate LLMs' cultural understanding, with a focus on Korean superstitions. The benchmark consists of 247 questions spanning 31 topics, assessing factual knowledge, culturally appropriate advice, and situational interpretation. We evaluate multilingual LLMs in both Korean and English to analyze their ability to reason about Korean cultural contexts and how language variations affect performance. To systematically assess cultural reasoning, we propose a novel evaluation strategy with customized scoring metrics that capture the extent to which models recognize cultural nuances and respond appropriately. Our findings highlight significant challenges in LLMs' cultural reasoning. While models generally recognize factual information, they struggle to apply it in practical scenarios. Furthermore, explicit cultural framing enhances performance more effectively than relying solely on the language of the prompt. To support further research, we publicly release Nunchi-Bench alongside a leaderboard.

文化推理韩语评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。