评估大模型与人类情感表达的对齐度,提升辅助沟通文本生成效果
Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
- 用人类判断评估大模型在情感表达上的语义对齐程度
- 词典式情感标签比数值化情绪维度更符合人类预期
- 不同情绪类型下,生成句子的情感感知受表达方式影响
语言模型在概念使用上与人类期望存在差距,尤其在通过辅助与替代沟通(AAC)工具帮助人们表达时尤为关键。本文提出‘表示对齐’评估任务,通过人工判断衡量这一差距。研究将关键词和情绪表示扩展为完整句子,采用四种情感表示:词语、词汇型与数值型的效价-唤醒-支配(VAD)维度,以及表情符号。除对齐度外,还评估了生成句的真实性和准确性。结果表明,当以英文词语(如“angry”)作为条件时,人们对大模型生成内容的认同度更高,显著优于数值型VAD表示;且不同情绪类型下,生成句的情感传达感知受表示形式影响显著。
原文摘要 · Abstract (English)
Gaps arise between a language model's use of concepts and people's expectations. This gap is critical when LLMs generate text to help people communicate via Augmentative and Alternative Communication (AAC) tools. In this work, we introduce the evaluation task of Representation Alignment for measuring this gap via human judgment. In our study, we expand keywords and emotion representations into full sentences. We select four emotion representations: Words, Valence-Arousal-Dominance (VAD) dimensions expressed in both Lexical and Numeric forms, and Emojis. In addition to Representation Alignment, we also measure people's judgments of the accuracy and realism of the generated sentences. While representations like VAD break emotions into easy-to-compute components, our findings show that people agree more with how LLMs generate when conditioned on English words (e.g., "angry") rather than VAD scales. This difference is especially visible when comparing Numeric VAD to words. Furthermore, we found that the perception of how much a generated sentence conveys an emotion is dependent on both the representation type and which emotion it is.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。