发现大模型在表达社交情绪时与人类文化规范存在系统性偏差。
Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms

- 构建心理启发的跨文化情绪表达评估框架
- 所有模型都更倾向表达亲社会情绪,尤其欧美人格表现异常
- 适合关注跨文化AI伦理与人机交互的研究者
表达具有社会功能的情绪(如彰显独立或促进互依)是人际互动的核心,且在不同文化中系统性差异。随着大模型越来越多地用于模拟跨文化情境下的行为,理解其是否忠实于人类的社会情绪表达模式至关重要。若模型输出与文化不匹配,其可用性将受损——尤其当用户误以为在与文化敏感的对话者互动时,可能采纳在特定文化中不恰当的建议。本文提出一种心理学启发的评估框架,通过对比欧裔美国人与拉丁美洲参与者对参与型与疏离型情绪的表达,评估六种前沿大模型在反映文化差异方面的表现。结果发现:所有模型均更倾向于表达参与型情绪,且在通常被良好建模的欧裔美国人角色上偏差尤为显著。进一步分析显示,模型响应高度集中且确定,未能捕捉人类在表达社会情绪时的多样性。消融实验表明这些模式对采样温度鲁棒,部分受提示语言影响,且依赖于响应生成格式。研究揭示当前大模型在文化与情感交互表征上的局限,尤其在表达社会情绪方面,对跨文化情感场景中的应用具有直接影响。
原文摘要 · Abstract (English)
The expression of emotions that serve social purposes, such as asserting independence or fostering interdependence, is central to human interactions and varies systematically across cultures. As LLMs are increasingly used to simulate human behavior in culturally nuanced interactions, it is important to understand whether they faithfully capture human patterns of social emotion expression. When LLM responses are not culturally aligned, their utility is compromised -- particularly when users assume they are interacting with a culturally attuned interlocutor, and may act on advice that proves inappropriate in their cultural context. We present a psychologically informed evaluation framework of cross-cultural social emotion expression in LLMs. Using a human study comparing European American and Latin American participants' expression of engaging and disengaging emotions, we evaluate six frontier LLMs on their ability to reflect culturally differentiated patterns for expressing social emotions. We find systematic misalignment between model and human behavior: all models express engaging emotions more than disengaging ones, with particularly stark differences observed for the generally well-represented European American persona. We further highlight that LLM responses are highly concentrated and deterministic, failing to capture the diversity of human responses in expressing social emotions. Our ablation analyses reveal that these patterns are robust to sampling temperatures, partially sensitive to prompt language, and dependent on the response elicitation format. Together, our findings highlight limitations in how current LLMs represent the interaction of cultural and emotional axes, particularly when expressing social emotions, with direct implications for their deployment in cross-cultural affective contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。