用无参数声学特征实现可解释的语音音色属性检测
Voice Timbre Attribute Detection with Compact and Interpretable Training-Free Acoustic Parameters
- 设计一组无训练参数的紧凑声学特征,捕捉音色关键声学量及其动态变化
- 在音色属性检测任务上超越传统倒谱特征和有监督神经网络嵌入
- 计算开销极低且物理意义明确,适合需要可解释性的语音分析场景
语音音色属性检测(vTAD)旨在判断不同语音语句间音色属性的相对强度。音色是语音感知中至关重要的成分,但其本质复杂。尽管深度神经网络(DNN)嵌入在说话人建模中表现优异,但常作为黑箱表示,缺乏物理可解释性且计算成本高。本文研究了一组紧凑的声学参数用于vTAD,该参数集捕捉了重要声学度量及其时序动态,被证实对任务至关重要。尽管结构简单,该参数集性能竞争力强,超越传统倒谱特征和有监督的DNN嵌入,并接近最先进的自监督模型。重要的是,所研究的参数集无需可训练参数,计算开销可忽略不计,且能为人类音色感知背后的物理特性提供清晰可解释性。
原文摘要 · Abstract (English)
Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural network (DNN) embeddings perform well in speaker modelling, they often act as black-box representations with limited physical interpretability and high computational cost. In this work, a compact acoustic parameter set is investigated for vTAD. The set captures important acoustic measures and their temporal dynamics which are found to be crucial in the task. Despite its simplicity, the acoustic parameter set is competitive, outperforming conventional cepstral features and supervised DNN embeddings, and approaching state-of-the-art self-supervised models. Importantly, the studied set require no trainable parameters, incur negligible computation, and offer explicit interpretability for analysing physical traits behind human timbre perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。