arXiv:2505.04382eess.AScs.LG2025-05被引 1

用离散最优传输提升语音转换质量与安全性

Discrete Optimal Transport and Voice Conversion

  • 基于预训练语音嵌入空间的离散最优传输框架
  • 显著改善语音分布对齐,降低语音识别错误率
  • 可伪装伪造语音以逃避检测,揭示安全风险

我们提出kDOT,一种在预训练语音嵌入空间中运行的离散最优传输(OT)框架,用于语音转换(VC)。与kNN-VC和SinkVC中的平均策略及MKL中的独立性假设不同,我们的方法利用离散OT计划的重心投影构建源与目标说话人嵌入分布间的传输映射。我们在LibriSpeech数据集上进行了全面消融实验,系统分析了传输嵌入数量及源/目标语句时长的影响。结果表明,采用重心投影的OT能持续改善分布对齐,在词错误率(WER)、主观评分(MOS)和音频差异度(FAD)上常优于基于平均的方法。此外,将离散OT作为后处理步骤可使伪造语音被最先进的伪造检测器误判为真实语音,展示了OT在嵌入空间中的强大域适应能力,同时也揭示了对伪造检测系统的重大安全隐忧。

原文摘要 · Abstract (English)

We propose kDOT, a discrete optimal transport (OT) framework for voice conversion (VC) operating in a pretrained speech embedding space. In contrast to the averaging strategies used in kNN-VC and SinkVC, and the independence assumption adopted in MKL, our method employs the barycentric projection of the discrete OT plan to construct a transport map between source and target speaker embedding distributions. We conduct a comprehensive ablation study over the number of transported embeddings and systematically analyze the impact of source and target utterance duration. Experiments on LibriSpeech demonstrate that OT with barycentric projection consistently improves distribution alignment and often outperforms averaging-based approaches in terms of WER, MOS, and FAD. Furthermore, we show that applying discrete OT as a post-processing step can transform spoofed speech into samples that are misclassified as bona fide by a state-of-the-art spoofing detector. This demonstrates the strong domain adaptation capability of OT in embedding space, while also revealing important security implications for spoof detection systems.

语音转换最优传输语音安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。