arXiv:2606.00407eess.AS2026-06ACL

让语音韵律表征不泄露身份信息,还能保持语音质量。

Privacy-preserving Prosody Representation Learning

论文配图:Privacy-preserving Prosody Representation Learning
图 1 · 摘自论文原文
  • 自监督学习中引入说话人解耦策略,分离韵律与身份特征。
  • 在音高重建和韵律事件检测任务上优于基线模型。
  • 适合注重隐私的语音生成与分析场景。

能捕捉语音韵律信息的表示对理解与生成均有益处,但声学韵律特征(如音高)会反映说话人身份,引发隐私泄露风险。为解决此问题,本文提出一种新的自监督方法,通过引入说话人解耦策略学习韵律表征。我们在三个任务上评估编码器的表征能力,包括音高重建与不同韵律事件检测。结果表明,该编码器在不损害韵律相关下游任务表现的前提下,显著提升了说话人解离效果,优于原始韵律特征与HuBERT-base基线。

原文摘要 · Abstract (English)

Speech representations that capture prosodic information can be useful for both understanding and generation. However, speaker characteristics are reflected in acoustic-prosodic features (e.g., pitch). To address privacy concerns from the leakage of identity information, we propose a new self-supervised approach to learning prosody representations that incorporates speaker disentanglement strategies. We evaluate our encoder on three tasks to probe representation capabilities, including pitch reconstruction and detection of different prosodic events. Our encoder outperforms raw prosody and HuBERT-base baselines, achieving strong speaker disentanglement without adverse impact on prosody-related downstream tasks.

语音隐私韵律表示自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。