通过排序对比学习,让语音和文本更好理解情绪的层次关系。
EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast
- 用情绪的高低、激活度排序来指导跨模态对比学习
- 在跨模态检索任务中优于现有方法,更准确捕捉情绪顺序
- 适合需要精细情绪理解的语音分析与人机交互场景
当前基于情感的对比语言-音频预训练(CLAP)方法通常简单地将音频样本与对应文本提示对齐,导致无法捕捉情绪的序数特性,削弱了情绪间的理解能力,并因对齐不足造成音视频嵌入之间的显著模态差距。为此,我们提出 EmotionRankCLAP,一种利用情绪语音的维度属性与自然语言提示的监督对比学习方法,联合建模细粒度情绪变化并增强跨模态对齐。该方法采用 Rank-N-Contrast 目标,基于情绪在效价-唤醒度空间中的排名进行样本对比,以学习有序关系。EmotionRankCLAP 在跨模态检索任务中表现出色,显著提升了多模态情绪序数建模能力。
原文摘要 · Abstract (English)
Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by naïvely aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal nature of emotions, hindering inter-emotion understanding and often resulting in a wide modality gap between the audio and text embeddings due to insufficient alignment. To handle these drawbacks, we introduce EmotionRankCLAP, a supervised contrastive learning approach that uses dimensional attributes of emotional speech and natural language prompts to jointly capture fine-grained emotion variations and improve cross-modal alignment. Our approach utilizes a Rank-N-Contrast objective to learn ordered relationships by contrasting samples based on their rankings in the valence-arousal space. EmotionRankCLAP outperforms existing emotion-CLAP methods in modeling emotion ordinality across modalities, measured via a cross-modal retrieval task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。