用混合语音关键帧增强短语音的说话人嵌入,提升检测效果
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
- 从混合语音中提取关键帧嵌入,加法融合增强原始嵌入
- 五次迭代后在短语音下达到全时长语音性能
- 适合低资源唤醒词场景的说话人活动检测应用
个人语音活动检测(PVAD)对于识别混合语音中的目标说话人片段至关重要,但其性能高度依赖说话人嵌入质量。一个关键实际限制是短注册语音(如唤醒词),提供的线索有限。本文提出一种自适应说话人嵌入自增强策略,通过将从混合语音中提取的关键帧嵌入与原始注册嵌入进行加法融合,增强嵌入表示。此外,引入长期适应策略,在检测过程中迭代优化嵌入,缓解说话人时间变化带来的影响。实验表明,在短注册条件下,召回率、精确率和F1分数均有显著提升,经过五次迭代更新后,性能可媲美全时长注册。源代码已公开于 https://anonymous.4open.science/r/ASE-PVAD-E5D6。
原文摘要 · Abstract (English)
Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings. A key practical limitation is the short enrollment speech--such as a wake-up word--which provides limited cues. This paper proposes a novel adaptive speaker embedding self-augmentation strategy that enhances PVAD performance by augmenting the original enrollment embeddings through additive fusion of keyframe embeddings extracted from mixed speech. Furthermore, we introduce a long-term adaptation strategy to iteratively refine embeddings during detection, mitigating speaker temporal variability. Experiments show significant gains in recall, precision, and F1-score under short enrollment conditions, matching full-length enrollment performance after five iterative updates. The source code is available at https://anonymous.4open.science/r/ASE-PVAD-E5D6 .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。