构建情感与风格对齐数据集,提升文本生成的情感细腻度与多样性。
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
- 基于细粒度情感分类,生成对话、正式、诗歌等多风格句子
- 通过困惑度、语义连贯性等指标验证情感真实性和语言多样性
- 适合研究情感控制、风格自适应生成和可解释性文本生成的学者
情感感知语言处理的进步正深刻影响对话AI、情感计算、计算心理学及创意内容生成等关键NLP应用。现有情感数据集或缺乏情感细微差别,或无法捕捉必要的风格多样性,限制了情感条件化文本生成系统的发展。为弥合这一关键缺口,本文提出一种系统构建的新数据集——ELSA(Emotion and Language Style Alignment Dataset),基于dair ai情感数据集和GoEmotions分类体系,采用细粒度情感分类。该数据集包含原始句子在对话、正式、诗歌、叙事等不同语境风格下的多个情感细腻变体,利用先进的大语言模型(LLMs)生成。通过困惑度、嵌入方差、可读性、词汇多样性及语义连贯性等指标进行严格计算评估,验证了数据集在情感真实性、语言流畅性与文本多样性方面的优势。综合指标分析证实其支持更深入探索情感条件化风格自适应文本生成的潜力。通过实现精细化情感调控的语言建模,本数据集为细粒度情感控制、提示驱动解释、可解释性及风格自适应表达性语言生成研究提供了坚实基础。
原文摘要 · Abstract (English)
Advancements in emotion aware language processing increasingly shape vital NLP applications ranging from conversational AI and affective computing to computational psychology and creative content generation. Existing emotion datasets either lack emotional granularity or fail to capture necessary stylistic diversity, limiting the advancement of effective emotion conditioned text generation systems. Seeking to bridge this crucial gap between granularity and style diversity, this paper introduces a novel systematically constructed dataset named ELSA Emotion and Language Style Alignment Dataset leveraging fine grained emotion taxonomies adapted from existing sources such as dair ai emotion dataset and GoEmotions taxonomy. This dataset comprises multiple emotionally nuanced variations of original sentences regenerated across distinct contextual styles such as conversational, formal, poetic, and narrative, using advanced Large Language Models LLMs. Rigorous computational evaluation using metrics such as perplexity, embedding variance, readability, lexical diversity, and semantic coherence measures validates the datasets emotional authenticity, linguistic fluency, and textual diversity. Comprehensive metric analyses affirm its potential to support deeper explorations into emotion conditioned style adaptive text generation. By enabling precision tuned emotionally nuanced language modeling, our dataset creates fertile ground for research on fine grained emotional control, prompt driven explanation, interpretability, and style adaptive expressive language generation with LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。