通过感知属性解析语音音色,实现音色差异的精准量化对比。
Introducing voice timbre attribute detection
- 基于说话人嵌入构建音色属性检测框架,用感知描述符比较语音强度。
- 在已见场景中ECAPA-TDNN表现更好,在未见场景中FACodec更具泛化能力。
- 适合语音感知研究与音色建模任务,开源数据集支持复现验证。
本文聚焦于解释语音信号所传达的音色特征,提出语音音色属性检测(vTAD)任务。该任务通过一组描述人类感知的感官属性来刻画音色。给定一对语音语句,系统在指定的音色描述符上比较其强度。提出的框架基于从语音中提取的说话人嵌入。在VCTK-RVA数据集上的实验表明:1)在测试说话人包含于训练集的已见场景中,ECAPA-TDNN说话人编码器表现更优;2)在测试说话人未出现在训练集的未见场景中,FACodec说话人编码器展现出更强的泛化能力。VCTK-RVA数据集与开源代码已公开于https://github.com/vTAD2025-Challenge/vTAD。
原文摘要 · Abstract (English)
This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human perception. A pair of speech utterances is processed, and their intensity is compared in a designated timbre descriptor. Moreover, a framework is proposed, which is built upon the speaker embeddings extracted from the speech utterances. The investigation is conducted on the VCTK-RVA dataset. Experimental examinations on the ECAPA-TDNN and FACodec speaker encoders demonstrated that: 1) the ECAPA-TDNN speaker encoder was more capable in the seen scenario, where the testing speakers were included in the training set; 2) the FACodec speaker encoder was superior in the unseen scenario, where the testing speakers were not part of the training, indicating enhanced generalization capability. The VCTK-RVA dataset and open-source code are available on the website https://github.com/vTAD2025-Challenge/vTAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。