首个跨文化情感理解基准,测试大模型在六种语言中的文化敏感度。
CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding
- 构建六语言文化情感问答数据集,每语言400题,需文化推理
- 发现不同文化对情绪理解差异大,模型表现随文化变化
- 英文提示加国家上下文比本地语言提示效果更好
自然语言处理研究日益关注情感分析等主观任务。然而,现有情感评测存在两大缺陷:(1) 多依赖关键词识别,忽略深层文化维度;(2) 多通过英语文本翻译生成,评估可靠性存疑。为此,我们提出文化情感视角(CuLEmo),首个跨六语言(阿姆哈拉语、阿拉伯语、英语、德语、印地语、西班牙语)的文化感知情感预测评测基准。每个语言包含400个精心设计的问题,需复杂文化推理。我们用该基准评估多个先进大模型在文化感知情感预测与情感分析任务上的表现。结果表明:(1) 情绪概念在不同语言文化间存在显著差异;(2) 大模型表现随语言与文化背景而异;(3) 使用英文提示并附加具体国家上下文,往往优于使用目标语言的提示。数据集与评测代码已公开。
原文摘要 · Abstract (English)
NLP research has increasingly focused on subjective tasks such as emotion analysis. However, existing emotion benchmarks suffer from two major shortcomings: (1) they largely rely on keyword-based emotion recognition, overlooking crucial cultural dimensions required for deeper emotion understanding, and (2) many are created by translating English-annotated data into other languages, leading to potentially unreliable evaluation. To address these issues, we introduce Cultural Lenses on Emotion (CuLEmo), the first benchmark designed to evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish. CuLEmo comprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. We use this benchmark to evaluate several state-of-the-art LLMs on culture-aware emotion prediction and sentiment analysis tasks. Our findings reveal that (1) emotion conceptualizations vary significantly across languages and cultures, (2) LLMs performance likewise varies by language and cultural context, and (3) prompting in English with explicit country context often outperforms in-language prompts for culture-aware emotion and sentiment understanding. The dataset and evaluation code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。