构建人格化情感数据集,让模型更懂不同人对同一事件的反应差异
Persona-E$^2$: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events

- 基于MBTI与大五人格标注真实人类情绪反应数据
- 现有大模型在社交媒体场景中难以准确捕捉情绪变化
- 人格特征显著提升模型理解力,缓解刻板印象问题
当前情感计算多将情绪视为文本的静态属性,聚焦作者情感而忽视读者视角。这种做法忽略了个体人格如何导致对同一事件的不同情绪评估。尽管角色扮演的大语言模型尝试模拟这些细微反应,但常陷入‘人格幻觉’——仅依赖表面刻板印象而非真实认知逻辑。关键瓶颈在于缺乏真实人类数据来建立人格特质与情绪转变之间的关联。为此,我们提出Persona-E$^2$(Persona-Event2Emotion),一个大规模数据集,基于标注的MBTI和大五人格特征,捕捉新闻、社交媒体及生活叙事中读者的情绪多样性。大量实验表明,当前顶尖大模型在捕捉精确情绪评估转变方面表现不佳,尤其在社交媒体领域。重要的是,我们发现人格信息能显著提升理解能力,其中大五人格特征有效缓解了‘人格幻觉’问题。
原文摘要 · Abstract (English)
Most affective computing research treats emotion as a static property of text, focusing on the writer's sentiment while overlooking the reader's perspective. This approach ignores how individual personalities lead to diverse emotional appraisals of the same event. Although role-playing Large Language Models (LLMs) attempt to simulate such nuanced reactions, they often suffer from "personality illusion'' -- relying on surface-level stereotypes rather than authentic cognitive logic. A critical bottleneck is the absence of ground-truth human data to link personality traits to emotional shifts. To bridge the gap, we introduce Persona-E$^2$ (Persona-Event2Emotion), a large-scale dataset grounded in annotated MBTI and Big Five traits to capture reader-based emotional variations across news, social media, and life narratives. Extensive experiments reveal that state-of-the-art LLMs struggle to capture precise appraisal shifts, particularly in social media domains. Crucially, we find that personality information significantly improves comprehension, with the Big Five traits alleviating "personality illusion.'
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。