利用预训练模型多层特征提升说话人识别准确率
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
- 直接处理预训练模型各层输出,提取固定长度语音向量
- 在多个数据集上误差率最低,性能优于现有方法
- 模型轻量高效,适合实际部署
自监督学习(SSL)在基于Transformer的模型上显著提升了说话人验证(SV)性能,提供了通用语音表征。然而,现有方法未能充分利用SSL编码器的多层结构特性。为此,我们提出层感知时延神经网络(L-TDNN),直接对预训练模型的逐层隐藏状态进行层/帧级处理,提取固定尺寸说话人向量。L-TDNN包含层感知卷积网络、帧自适应层聚合与注意力统计池化,显式建模被忽视的层维度。我们在多个语音SSL Transformer和多样化语音-说话人语料库上评估了L-TDNN,并与其它利用预训练编码器的方法对比。实验表明,L-TDNN始终表现出稳健的验证性能,误差率最低。同时,在模型紧凑性和推理效率方面也达到现有系统水平。结果凸显了所提层感知处理方法的优势。未来工作包括探索与SSL前端联合训练及引入分数校准以进一步提升性能。
原文摘要 · Abstract (English)
Recent advances in self-supervised learning (SSL) on Transformers have significantly improved speaker verification (SV) by providing domain-general speech representations. However, existing approaches have underutilized the multi-layered nature of SSL encoders. To address this limitation, we propose the layer-aware time-delay neural network (L-TDNN), which directly performs layer/frame-wise processing on the layer-wise hidden state outputs from pre-trained models, extracting fixed-size speaker vectors. L-TDNN comprises a layer-aware convolutional network, a frame-adaptive layer aggregation, and attentive statistic pooling, explicitly modeling of the recognition and processing of previously overlooked layer dimension. We evaluated L-TDNN across multiple speech SSL Transformers and diverse speech-speaker corpora against other approaches for leveraging pre-trained encoders. L-TDNN consistently demonstrated robust verification performance, achieving the lowest error rates throughout the experiments. Concurrently, it stood out in terms of model compactness and exhibited inference efficiency comparable to the existing systems. These results highlight the advantages derived from the proposed layer-aware processing approach. Future work includes exploring joint training with SSL frontends and the incorporation of score calibration to further enhance state-of-the-art verification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。