arXiv:2601.12591cs.SDeess.AS2026-01中稿 · ICASSP 2026被引 2

用软标签提升情感音频文本模型的模糊边界识别能力

SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing

  • 引入模态内相似性和副语言特征生成软目标
  • 在8个跨语言任务上均优于传统CLAP方法
  • 适合需要细粒度情感理解的语音应用

人类情感的模糊性给机器学习模型带来挑战,因情感之间常重叠且边界不清。对比语言-音频预训练(CLAP)已成为通用情感识别的关键技术。然而,传统CLAP强制音频-文本样本一一对应,忽略模态内相似性,将所有不匹配对视为同等负样本,与情感间模糊边界矛盾。为此,本文提出SmoothCLAP,利用模态内相似性和副语言特征生成软目标,结合传统对比监督,使模型学习到反映情感梯度关系的嵌入表示,同时保持与CLAP相同的推理流程。在英语和德语的八个情感计算任务上的实验表明,SmoothCLAP始终表现更优。结果表明,利用软监督是构建情感感知音频-文本模型的可行策略。

原文摘要 · Abstract (English)

The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretraining (CLAP) has emerged as a key technique for generalisable emotion recognition. However, as conventional CLAP enforces a strict one-to-one alignment between paired audio-text samples, it overlooks intra-modal similarity and treats all non-matching pairs as equally negative. This conflicts with the fuzzy boundaries between different emotions. To address this limitation, we propose SmoothCLAP, which introduces softened targets derived from intra-modal similarity and paralinguistic features. By combining these softened targets with conventional contrastive supervision, SmoothCLAP learns embeddings that respect graded emotional relationships, while retaining the same inference pipeline as CLAP. Experiments on eight affective computing tasks across English and German demonstrate that SmoothCLAP is consistently achieving superior performance. Our results highlight that leveraging soft supervision is a promising strategy for building emotion-aware audio-text models.

情感计算音频文本对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。