用数学算子分离语音中的说话人特征与内容,无需文本标注。
Koopman Regularized Deep Speech Disentanglement for Speaker Verification
- 引入科普曼算子建模语音动态,结合实例归一化实现特征解耦。
- 在多个数据集上达到顶尖验证性能,内容错误率保持高位。
- 参数少、不依赖文本监督,适合实际部署和长期维护。
人类语音同时包含语言内容和说话人特有特征,使说话人验证成为身份关键应用的核心技术。现代深度学习说话人验证系统旨在学习对语义内容和环境噪声等干扰因素不变的说话人表征。然而,许多现有方法依赖标签数据、文本监督或大型预训练模型作为特征提取器,限制了可扩展性和实际部署,引发可持续性担忧。我们提出深度科普曼语音解耦自编码器(DKSD-AE),一种结构化自编码器,结合新颖的多步科普曼算子学习模块与实例归一化,以解耦说话人与内容动态。跨多个数据集的定量实验表明,DKSD-AE 在性能上优于或媲美最先进基线,同时保持高内容错误率,证实了有效解耦。该结果在显著更少参数下达成,且无需文本监督。此外,性能在评估规模扩大时仍稳定,凸显表征鲁棒性与泛化能力。研究结果表明,结合实例归一化的科普曼时序建模,为聚焦说话人的表征学习提供了高效而原则性的解决方案。
原文摘要 · Abstract (English)
Human speech contains both linguistic content and speaker dependent characteristics making speaker verification a key technology in identity critical applications. Modern deep learning speaker verification systems aim to learn speaker representations that are invariant to semantic content and nuisance factors such as ambient noise. However, many existing approaches depend on labelled data, textual supervision or large pretrained models as feature extractors, limiting scalability and practical deployment, raising sustainability concerns. We propose Deep Koopman Speech Disentanglement Autoencoder (DKSD-AE), a structured autoencoder that combines a novel multi-step Koopman operator learning module with instance normalization to disentangle speaker and content dynamics. Quantitative experiments across multiple datasets demonstrate that DKSD-AE achieves improved or competitive speaker verification performance compared to state-of-the-art baselines while maintaining high content EER, confirming effective disentanglement. These results are obtained with substantially fewer parameters and without textual supervision. Moreover, performance remains stable under increased evaluation scale, highlighting representation robustness and generalization. Our findings suggest that Koopman-based temporal modelling, when combined with instance normalization, provides an efficient and principled solution for speaker-focused representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。