让语音模型直接理解情绪与人格,无需后续文本模型
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
- 用语言模型做教师,指导语音模型对齐语义与心理特征
- 在50万段心理访谈音频上训练,心理任务误差降低83.8%
- 适合心理健康分析、情感计算等需要深层语义理解的场景
当前语音编码常依赖额外的文本语言模型以获得稳健的人类交流表征,尽管最先进的语音转文本模型内部已包含语言模型。本文提出WhiSPA(Whisper带语义与心理对齐),通过引入新颖的音频训练目标:以语言模型嵌入作为教师的对比损失。基于超过50万段心理健康音频访谈数据,评估了将Whisper隐空间与文本自编码器(SBERT)的语义表示及情绪、人格等基本心理维度的词法嵌入对齐的效果。在自监督情感任务和下游心理任务中,WhiSPA均超越现有语音编码器,平均误差分别降低73.4%和83.8%。结果表明,无需在语音转文本输出后运行后续文本语言模型,即可获得丰富的人类交流心理表征。
原文摘要 · Abstract (English)
Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the LM within an audio model such that the subsequent text-LM is unnecessary. We introduce WhiSPA (Whisper with Semantic and Psychological Alignment), which leverages a novel audio training objective: contrastive loss with a language model embedding as a teacher. Using over 500k speech segments from mental health audio interviews, we evaluate the utility of aligning Whisper's latent space with semantic representations from a text autoencoder (SBERT) and lexically derived embeddings of basic psychological dimensions: emotion and personality. Over self-supervised affective tasks and downstream psychological tasks, WhiSPA surpasses current speech encoders, achieving an average error reduction of 73.4% and 83.8%, respectively. WhiSPA demonstrates that it is not always necessary to run a subsequent text LM on speech-to-text output in order to get a rich psychological representation of human communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。