用语音对比预测听众印象差异,自监督模型表现更优。
Investigation for Relative Voice Impression Estimation
- 通过对比同一人不同语调的录音,预测印象变化方向
- 自监督模型在复杂印象(如冷暖)上显著优于传统特征
- 适合研究语音情感、声音设计与人机交互的学者
语音的副语言和非语言特征强烈影响听者印象。现有研究多关注绝对印象评分,本文首次系统探究相对语音印象估计(RIE),即从同一说话人两个语句中预测其感知差异。目标为基于主观评价生成的低维向量,量化第二个语句相对于第一个在反义轴(如“暗—亮”)上的感知偏移。实验采用专业朗读者以多种风格朗读同一文本的录音,对比三类建模方法:经典声学特征、自监督语音表征和多模态大语言模型(MLLM)。结果表明,自监督语音模型在捕捉复杂动态印象(如“冷—暖”)方面显著优于传统声学特征,后者在该任务中表现不佳;而当前的MLLM在细粒度成对任务中不可靠。本研究首次系统验证了RIE框架,并证实自监督语音模型在捕捉细微感知变化中的优势。
原文摘要 · Abstract (English)
Paralinguistic and non-linguistic aspects of speech strongly influence listener impressions. While most research focuses on absolute impression scoring, this study investigates relative voice impression estimation (RIE), a framework for predicting the perceptual difference between two utterances from the same speaker. The estimation target is a low-dimensional vector derived from subjective evaluations, quantifying the perceptual shift of the second utterance relative to the first along an antonymic axis (e.g., ``Dark--Bright''). To isolate expressive and prosodic variation, we used recordings of a professional speaker reading a text in various styles. We compare three modeling approaches: classical acoustic features commonly used for speech emotion recognition, self-supervised speech representations, and multimodal large language models (MLLMs). Our results demonstrate that models using self-supervised representations outperform methods with classical acoustic features, particularly in capturing complex and dynamic impressions (e.g., ``Cold--Warm'') where classical features fail. In contrast, current MLLMs prove unreliable for this fine-grained pairwise task. This study provides the first systematic investigation of RIE and demonstrates the strength of self-supervised speech models in capturing subtle perceptual variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。