arXiv:2510.05291cs.CL2025-10中稿 · EMNLP被引 2

测试大模型在九种亚洲语言中的文化偏见,发现模型对东西方文化存在显著偏好差异。

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

  • 构建涵盖六种亚洲文化的19,530个实体的基准测试集Camellia。
  • 多模型测试显示,不同地区开发的模型在文化适配上表现差异显著。
  • 模型在亚洲语言中对文化语境理解不足,导致实体提取能力下降。

随着大语言模型(LLMs)多语言能力增强,其对文化多样性实体的敏感性日益重要。先前研究(Naous et al., 2024)表明,LLMs在阿拉伯语中倾向于偏好西方相关实体。由于缺乏以实体为中心的多语言基准,非西方语言中是否存在类似偏见尚不明确。本文提出Camellia,一个用于评估九种亚洲语言中实体导向文化偏见的基准,覆盖六种亚洲文化。Camellia包含19,530个经人工标注的、与亚洲或西方文化相关的实体,以及从社交媒体帖子中提取的2,173个掩码上下文。我们利用Camellia在三个任务上评估四种近期多语言LLMs的文化偏见:文化上下文适应、情感关联和实体抽取问答。分析显示,这些模型在跨语言文化适配上表现不佳,性能随模型开发区域而异。不同模型家族表现出独特偏见,体现在其将文化与特定情感关联的方式上。此外,模型在部分亚洲语言中存在上下文理解困难,导致文化间实体提取性能差距。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly important. Prior work by Naous et al. (2024) has shown that LLMs often favor Western-associated entities in Arabic. Due to the lack of entity-centric multilingual benchmarks, it remains unclear if such biases also manifest in various non-Western languages. In this paper, we introduce Camellia, a benchmark for evaluating entity-centric cultural biases in nine Asian languages, spanning six Asian cultures. Camellia includes 19,530 manually annotated entities associated with the covered Asian or Western cultures, as well as 2,173 masked contexts for these entities derived from social media posts. Using Camellia, we evaluate cultural biases in four recent multilingual LLMs across three tasks: cultural context adaptation, sentiment association, and entity extractive QA. Our analyses show that LLMs struggle with cultural adaptation across these languages, with performance differing across models developed in different regions. We further observe that different LLM families can hold distinct biases, reflected in the ways they link cultures to particular sentiments. Lastly, we find that LLMs can struggle with context understanding in some Asian languages, creating performance gaps between cultures in entity extraction.

文化偏见多语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。