arXiv:2606.16115eess.AScs.SD2026-06中稿 · Interspeech 2026

用混合注册+神经重打分,提升短语音说话人识别稳定性

Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment

论文配图:Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment
图 1 · 摘自论文原文
  • 结合文本依赖与文本无关注册,通过并行交叉注意力进行帧级比对
  • 在VoxPhrase数据集上,多模型表现持续提升,最长达12%相对增益
  • 适合需要高鲁棒性短语音验证的场景,如语音助手唤醒

短时说话人验证(SDSV)对个性化关键词检测至关重要,测试语句通常不足三秒。语音时长受限导致说话人表征不稳定,对噪声和发音变化更敏感,性能下降。为此,我们从VoxCeleb数据集自动构建了大规模的SDSV语料库VoxPhrase。分析表明,文本依赖(TD)注册受时长限制,表征不稳定;而文本无关(TI)注册虽存在内容不匹配,但随着注册时长增加,表征更稳定。为此,我们提出一种混合注册神经重打分框架,融合TD与TI注册,并通过并行交叉注意力实现帧级比较。在VoxPhrase上的实验显示,多个说话人模型均获得一致性能提升。

原文摘要 · Abstract (English)

Short-duration speaker verification (SDSV) is crucial for personalized keyword spotting, where test utterances are typically shorter than three seconds. Limited speech duration results in unstable speaker representations and increased sensitivity to noise and phoneme variations, thereby degrading performance. To investigate this issue, we construct VoxPhrase, a large-scale SDSV corpus automatically segmented from the VoxCeleb dataset. Our analysis shows that text-dependent (TD) enrollment is constrained by duration and yields unstable speaker representations. In contrast, although text-independent (TI) enrollment introduces content mismatch, its representations become more stable as the enrollment duration increases. Accordingly, we propose a hybrid-enrollment neural re-scoring framework that combines TD and TI enrollment and performs frame-level comparison via parallel cross-attention. Experiments on VoxPhrase demonstrate consistent improvements across multiple speaker models.

说话人验证短语音神经重打分混合注册

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。