arXiv:2410.09636eess.AScs.AI2024-10中稿 · APSIPA 2024 ASC

用零样本语音情绪识别直接估计购买意愿,效果媲美有监督模型。

Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?

  • 基于对比语言音频预训练框架,支持多任务多类别零样本情绪识别。
  • 在未知双极情绪上表现优异,零样本估计效果接近有监督模型。
  • 首次实现从语音直接预测购买意愿,适合电商和用户行为研究者。

本文提出一种零样本语音情绪识别(SER)方法,可识别未在模型训练中定义的情绪。传统方法仅限于单个词语定义的情绪识别,而本研究旨在识别如“我想买 - 我不想买”这类未知的双极情绪。为此,方法在对比语言-音频预训练(CLAP)框架基础上,引入多类别与多任务设置,使模型能自由使用句子定义类别,并评估未知双极情绪。研究聚焦购买意愿这一双极情绪,探究模型在零样本条件下的表现。这是首次直接从语音中估计购买意愿的研究。实验表明,该方法的零样本估计结果达到与有监督学习模型相当的水平。

原文摘要 · Abstract (English)

This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a single word. Moreover, we have the motivation to recognize unknown bipolar emotions such as ``I want to buy - I do not want to buy.'' In order to allow the model to define classes using sentences freely and to estimate unknown bipolar emotions, our proposed method expands upon the contrastive language-audio pre-training (CLAP) framework by introducing multi-class and multi-task settings. We also focus on purchase intention as a bipolar emotion and investigate the model's performance to zero-shot estimate it. This study is the first attempt to estimate purchase intention from speech directly. Experiments confirm that the results of zero-shot estimation by the proposed method are at the same level as those of the model trained by supervised learning.

语音识别零样本购买意愿情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。