用距离感知软提示提升多模态情绪连续估计效果
Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation
- 将情绪空间分区域,用高斯核计算软标签以捕捉细微情感变化
- 在Aff-Wild2数据集上优于官方基线,对真实场景数据鲁棒
- 适合关注情绪识别、多模态融合与连续标注的研究者
情绪的效价-唤醒度(VA)估计对于捕捉自然环境中人类情感的细微差别至关重要。尽管预训练的视觉语言模型(如CLIP)具备出色的语义对齐能力,但其在连续回归任务中的应用常受限于文本提示的离散性。本文提出一种新型多模态VA估计框架,引入距离感知软提示引导机制,弥合语义表征与连续情感维度之间的差距。具体而言,我们将VA空间划分为多个离散区域,每个区域对应不同的文本描述。不采用硬分类,而是利用高斯核根据真实坐标与区域中心的欧氏距离计算软标签,使模型能够学习细粒度的情感过渡。多模态融合方面,架构采用CLIP图像编码器和音频频谱变换器提取鲁棒的视觉与声学特征,通过门控循环单元进行时序建模,并通过分层融合策略依次结合跨模态注意力对齐与门控融合自适应优化。在Aff-Wild2数据集上的实验表明,所提语义引导方法优于官方基线,并在真实场景数据上表现出强鲁棒性。
原文摘要 · Abstract (English)
Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision-language models such as CLIP have demonstrated remarkable semantic alignment capabilities, their application to continuous regression tasks is often limited by the discrete nature of text prompts. In this paper, we propose a novel multimodal framework for VA estimation that introduces Distance-aware Soft Prompt Guidance to bridge the gap between semantic representations and continuous affective dimensions. Specifically, we partition the VA space into multiple discrete regions, each associated with distinct textual descriptions. Rather than relying on hard categorization, we employ a Gaussian kernel to compute soft labels based on the Euclidean distance between the ground-truth coordinates and the region centers, allowing the model to learn fine-grained emotional transitions. For multimodal integration, our architecture utilizes a CLIP image encoder and an Audio Spectrogram Transformer to extract robust visual and acoustic features. These features are temporally modeled using Gated Recurrent Units and integrated through a hierarchical fusion scheme that sequentially combines cross-modal attention for alignment and gated fusion for adaptive refinement. Experimental results on the Aff-Wild2 dataset show that the proposed semantic-guided approach outperforms the official baseline and demonstrates robust performance on in-the-wild data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。