arXiv:2410.15499cs.AIcs.SD2024-10中稿 · NeurIPS被引 2

用符合人耳感知的损失函数提升语音匿名化音质

Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses

  • 设计基于人类听觉系统的感知损失,不依赖特定模型架构
  • 在多个数据集与语种上显著提升语音自然度和可懂度
  • 适合语音匿名化、隐私保护等注重音质的应用场景

云语音助手的普及加剧了对有效语音匿名化的需求,旨在隐藏说话人身份的同时保留后续任务所需的关键信息。一种实现方式是语音转换。现有方法多关注复杂架构与训练技巧,而本研究强调源于人类听觉系统的损失函数的重要性。所提出的损失函数具有模型无关性,结合手工设计与深度学习特征,有效捕捉语音质量表征。通过客观与主观评估,我们证明:采用感知驱动损失的VQVAE模型,在自然度、可懂度和语调保持方面均优于基线模型,且在多种数据集、语言、目标说话人及性别下表现一致稳定。

原文摘要 · Abstract (English)

The increasing use of cloud-based speech assistants has heightened the need for effective speech anonymization, which aims to obscure a speaker's identity while retaining critical information for subsequent tasks. One approach to achieving this is through voice conversion. While existing methods often emphasize complex architectures and training techniques, our research underscores the importance of loss functions inspired by the human auditory system. Our proposed loss functions are model-agnostic, incorporating handcrafted and deep learning-based features to effectively capture quality representations. Through objective and subjective evaluations, we demonstrate that a VQVAE-based model, enhanced with our perception-driven losses, surpasses the vanilla model in terms of naturalness, intelligibility, and prosody while maintaining speaker anonymity. These improvements are consistently observed across various datasets, languages, target speakers, and genders.

语音匿名化感知损失语音转换隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。