arXiv:2602.14172cs.SDcs.CL2026-02中稿 · Speech Prosody 202…被引 1

用语音对比预测听众印象差异,自监督模型表现更优。

Investigation for Relative Voice Impression Estimation

  • 通过对比同一人不同语调的录音,预测印象变化方向
  • 自监督模型在复杂印象(如冷暖)上显著优于传统特征
  • 适合研究语音情感、声音设计与人机交互的学者

语音的副语言和非语言特征强烈影响听者印象。现有研究多关注绝对印象评分,本文首次系统探究相对语音印象估计(RIE),即从同一说话人两个语句中预测其感知差异。目标为基于主观评价生成的低维向量,量化第二个语句相对于第一个在反义轴(如“暗—亮”)上的感知偏移。实验采用专业朗读者以多种风格朗读同一文本的录音,对比三类建模方法:经典声学特征、自监督语音表征和多模态大语言模型(MLLM)。结果表明,自监督语音模型在捕捉复杂动态印象(如“冷—暖”)方面显著优于传统声学特征,后者在该任务中表现不佳;而当前的MLLM在细粒度成对任务中不可靠。本研究首次系统验证了RIE框架,并证实自监督语音模型在捕捉细微感知变化中的优势。

原文摘要 · Abstract (English)

Paralinguistic and non-linguistic aspects of speech strongly influence listener impressions. While most research focuses on absolute impression scoring, this study investigates relative voice impression estimation (RIE), a framework for predicting the perceptual difference between two utterances from the same speaker. The estimation target is a low-dimensional vector derived from subjective evaluations, quantifying the perceptual shift of the second utterance relative to the first along an antonymic axis (e.g., ``Dark--Bright''). To isolate expressive and prosodic variation, we used recordings of a professional speaker reading a text in various styles. We compare three modeling approaches: classical acoustic features commonly used for speech emotion recognition, self-supervised speech representations, and multimodal large language models (MLLMs). Our results demonstrate that models using self-supervised representations outperform methods with classical acoustic features, particularly in capturing complex and dynamic impressions (e.g., ``Cold--Warm'') where classical features fail. In contrast, current MLLMs prove unreliable for this fine-grained pairwise task. This study provides the first systematic investigation of RIE and demonstrates the strength of self-supervised speech models in capturing subtle perceptual variations.

语音印象自监督学习感知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。