arXiv:2509.19721eess.AS2025-09中稿 · ICASSP 2026

用多分辨率时域编码器提升短段语音验证精度

Short-Segment Speaker Verification with Pre-trained Models and Multi-Resolution Encoder

  • 融合预训练模型、滤波器组与多分辨率时域特征
  • 在1.56-12.5毫秒窗口下实现更精细的语音表征
  • 适合低时长语音验证,尤其对<2秒片段有效

利用自监督学习预训练模型提取的声纹验证(SV)特征近期表现优异。然而,这些预训练模型(PTM)通常具有20毫秒的时间分辨率,低于常规滤波器组特征。这在输入段落短于2秒的短段语音验证中尤为不利,因需从有限长度中提取尽可能多的信息。尽管已有研究尝试使用HuBERT模型的多分辨率特征,但采样率为16kHz时的窗口偏移仅为20、40和100毫秒,仅考虑了较低分辨率特征。本研究提出一种新声纹验证系统,结合预训练模型特征、滤波器组特征以及来自多分辨率时域编码器的特征,其窗口偏移分别为1.56、3.13、6.25和12.5毫秒。在VoxCeleb数据集上,针对不同输入长度的实验结果表明,该系统在多种特征组合下均表现出一致的性能提升。

原文摘要 · Abstract (English)

Speaker verification (SV) utilizing features obtained from models pre-trained via self-supervised learning has recently demonstrated impressive performances. However, these pre-trained models (PTMs) usually have a temporal resolution of 20 ms, which is lower than typical filterbank features. It may be problematic especially for short-segment SV with an input segment shorter than 2 s, in which we need to extract as much information as possible from the input with a limited length. Although there have been approaches to utilize multi-resolution features from the HuBERT models, the window shifts were 20, 40, and 100 ms when the sampling rate was 16 kHz and thus only lower resolution features were considered. In this study, we propose an SV system which utilizes PTM features along with filterbank features and those from the multi-resolution time domain encoder with window shifts of 1.56, 3.13, 6.25, and 12.5 ms. Experimental results on the VoxCeleb dataset with various input lengths showed consistent improvements over systems with various combinations of input features.

声纹验证多分辨率预训练模型短语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。