用强化学习生成更真实的情感描述,提升语音情感 caption 的准确与多样。
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
- 基于情绪感知的强化学习策略,动态优化生成过程。
- 在 EmotionTalk 数据集上准确率与多样性显著提升。
- 适合需要自然语言情感表达的研究者和开发者。
语音情感描述(SEC)已成为重要研究方向。人类语音中情感内容的复杂性使传统离散分类方法难以充分表征。因此,使用自然语言描述语音情感成为更有效捕捉和表达情感的新途径。本文提出 MECap-R1,一种面向多模态情感描述的新型情绪感知强化学习策略。通过采用情绪感知的组相对策略优化(Emo-GRPO),该框架能精准捕捉情感与语义特征,克服了固定规则在处理描述动态性和灵活性方面的不足。在 EmotionTalk 数据集上的实验表明,MECap-R1 在生成情感描述方面表现优异,准确率与多样性均取得显著提升。
原文摘要 · Abstract (English)
Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discrete classification methods to provide an adequate representation. Consequently, utilizing natural language to describe speech emotions presents a novel avenue for more effectively capturing and expressing affect. In this paper, we propose MECap-R1, a pioneering emotion-aware policy with reinforcement learning for multimodal emotion captioning. By employing Group Relative Policy Optimization with emotion-aware reward (Emo-GRPO), the framework precisely captures the emotion and semantic features, thereby addressing the shortcomings of rigid rules in handling the dynamic and flexible nature of captions. Experimental results on the EmotionTalk dataset demonstrate that MECap-R1 performs well in generating emotion descriptions and achieves substantial gains in both accuracy and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。