用多阶段人格条件生成情感文本,解决情绪识别数据稀缺问题
Persona-Based Synthetic Data Generation Using Multi-Stage Conditioning with Large Language Models for Emotion Recognition
- 通过人格属性与情境构建虚拟角色,分阶段引导大模型生成情绪表达
- 生成数据在语义多样性、真实感和分类任务表现上均优于基线方法
- 适合需要高质量情绪数据的NLP研究者,尤其关注个性化情感建模
在情绪识别领域,高性能模型的发展受限于高质量、多样化的情绪数据集稀缺。情绪表达具有主观性,受个体人格特质、社会文化背景和情境因素影响,大规模可泛化数据收集在伦理与实践上均具挑战。为此,我们提出PersonaGen框架,利用大语言模型(LLM)通过多阶段人格条件生成丰富情绪文本。PersonaGen结合人口统计特征、社会文化背景与具体情境,构建分层虚拟人格,用于指导情绪表达生成。我们对合成数据进行全面评估:通过聚类与分布度量检验语义多样性;采用基于LLM的质量评分衡量人类相似性;通过与真实情绪语料对比验证现实感;并在下游情绪分类任务中测试实用性。实验表明,PersonaGen在生成多样化、连贯且具备区分性的表情达意方面显著优于基线方法,展现出作为真实情绪数据集增强或替代方案的强大潜力。
原文摘要 · Abstract (English)
In the field of emotion recognition, the development of high-performance models remains a challenge due to the scarcity of high-quality, diverse emotional datasets. Emotional expressions are inherently subjective, shaped by individual personality traits, socio-cultural backgrounds, and contextual factors, making large-scale, generalizable data collection both ethically and practically difficult. To address this issue, we introduce PersonaGen, a novel framework for generating emotionally rich text using a Large Language Model (LLM) through multi-stage persona-based conditioning. PersonaGen constructs layered virtual personas by combining demographic attributes, socio-cultural backgrounds, and detailed situational contexts, which are then used to guide emotion expression generation. We conduct comprehensive evaluations of the generated synthetic data, assessing semantic diversity through clustering and distributional metrics, human-likeness via LLM-based quality scoring, realism through comparison with real-world emotion corpora, and practical utility in downstream emotion classification tasks. Experimental results show that PersonaGen significantly outperforms baseline methods in generating diverse, coherent, and discriminative emotion expressions, demonstrating its potential as a robust alternative for augmenting or replacing real-world emotional datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。