首个跨文化人格评估基准,让大模型学会说不同文化的话。
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
- 构建6国3000个情境问答,测试模型在本地价值观下的性格表达
- 相比现有基准,人格分布偏差降低超20%(Wasserstein距离)
- 适合研究跨文化AI、人机交互与全球化模型设计的团队
随着大语言模型广泛应用于教学、心理健康等交互场景,其在不同文化背景下以恰当方式表达个性的能力愈发重要。现有研究多忽略文化与人格的相互作用。为此,我们提出CulturalPersonas,首个基于人类验证的大规模基准,用于评估模型在文化语境中的个性表达。该数据集涵盖六个国家的3000个情景化问题,聚焦本土价值观下的日常行为。我们评估了三种大语言模型,采用多选与开放回答两种形式。结果表明,CulturalPersonas使模型输出更贴近各国真实人格分布(跨模型和国家平均降低超过20%的Wasserstein距离),且生成内容更具表现力与文化一致性。该基准能揭示模型对文化提示的有差异响应,为对齐全球行为规范提供新路径。通过融合人格与文化细微差别,CulturalPersonas有望推动更具备社会智能与全球适应性的大模型发展。
原文摘要 · Abstract (English)
As LLMs become central to interactive applications, ranging from tutoring to mental health, the ability to express personality in culturally appropriate ways is increasingly important. While recent works have explored personality evaluation of LLMs, they largely overlook the interplay between culture and personality. To address this, we introduce CulturalPersonas, the first large-scale benchmark with human validation for evaluating LLMs' personality expression in culturally grounded, behaviorally rich contexts. Our dataset spans 3,000 scenario-based questions across six diverse countries, designed to elicit personality through everyday scenarios rooted in local values. We evaluate three LLMs, using both multiple-choice and open-ended response formats. Our results show that CulturalPersonas improves alignment with country-specific human personality distributions (over a 20% reduction in Wasserstein distance across models and countries) and elicits more expressive, culturally coherent outputs compared to existing benchmarks. CulturalPersonas surfaces meaningful modulated trait outputs in response to culturally grounded prompts, offering new directions for aligning LLMs to global norms of behavior. By bridging personality expression and cultural nuance, we envision that CulturalPersonas will pave the way for more socially intelligent and globally adaptive LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。