通过差异注意力建模语音音色属性,提升跨说话人识别效果。
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
- 基于成对比较与差异注意力,捕捉音色属性的相对差异。
- 在VCTK-RVA数据集上显著提升多属性检测性能,跨说话人泛化能力增强。
- 适合语音合成、音色编辑等需要精细音色控制的研究者。
语音音色属性检测(vTAD)在语音生成任务中的细粒度音色建模中至关重要。然而,由于音色描述的主观性以及现有数据集中的严重标签不平衡,该任务仍具挑战。本文提出QvTAD,一种基于差异注意力的成对比较框架,以增强感知音色属性的建模。为解决VCTK-RVA数据集中的标签不平衡问题,我们引入一种基于图的数据增强策略,构建有向无环图,并使用并查集技术自动挖掘未观测到的、具有有效属性比较的语句对。该框架利用预训练FACodec提取的说话人嵌入,并引入相对音色偏移感知的差异注意力模块,通过差异去噪与对比增强机制,显式建模成对语句间的属性特异性差异。在VCTK-RVA基准上的实验结果表明,QvTAD在多个音色描述符上均取得显著提升,尤其在跨说话人泛化场景下表现突出。
原文摘要 · Abstract (English)
Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subjective nature of timbre descriptors and the severe label imbalance in existing datasets. In this work, we present QvTAD, a novel pairwise comparison framework based on differential attention, designed to enhance the modeling of perceptual timbre attributes. To address the label imbalance in the VCTK-RVA dataset, we introduce a graph-based data augmentation strategy that constructs a Directed Acyclic Graph and employs Disjoint-Set Union techniques to automatically mine unobserved utterance pairs with valid attribute comparisons. Our framework leverages speaker embeddings from a pretrained FACodec, and incorporates a Relative Timbre Shift-Aware Differential Attention module. This module explicitly models attribute-specific contrasts between paired utterances via differential denoising and contrast amplification mechanisms. Experimental results on the VCTK-RVA benchmark demonstrate that QvTAD achieves substantial improvements across multiple timbre descriptors, with particularly notable gains in cross-speaker generalization scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。