arXiv:2505.21809cs.SDcs.LG2025-05中稿 · Interspeech 2025被引 5

用可解释的语音质量维度建模异常发音与情感表达。

Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect

  • 基于预训练模型嵌入,训练七维语音质量探测器。
  • 在多个数据集上表现优异,零样本迁移能力强。
  • 适合语音康复、情感分析等需要可解释性的场景。

语音质量感知维度能刻画异常发音及其他语音变化特征。本文针对七种语音维度(可理解性、辅音不清晰、声音粗糙、自然度、单调响度、单调音高和气息声)构建并评估了语音质量模型。模型在包含11,184个样本、来自434名说话者的公开语音可访问性(SAP)数据集上训练,采用冻结预训练模型的嵌入作为特征。结果显示,探测器在SAP数据集的不同语音诱发类别中均具备强性能与强泛化能力。进一步在意大利语异常发音、英语异常发音及情感语音等额外数据集上验证了零样本性能,结果表明其在未见语言和任务中仍表现良好。该研究证实,语音质量维度在语音风格相关任务中具有实用价值。

原文摘要 · Abstract (English)

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise consonants, harsh voice, naturalness, monoloudness, monopitch, and breathiness). Probes were trained on the public Speech Accessibility (SAP) project dataset with 11,184 samples from 434 speakers, using embeddings from frozen pre-trained models as features. We found that our probes had both strong performance and strong generalization across speech elicitation categories in the SAP dataset. We further validated zero-shot performance on additional datasets, encompassing unseen languages and tasks: Italian atypical speech, English atypical speech, and affective speech. The strong zero-shot performance and the interpretability of results across an array of evaluations suggests the utility of using voice quality dimensions in speaking style-related tasks.

语音质量异常发音零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。