arXiv:2512.06304eess.AScs.AI2025-12综述被引 1

研究语音转换模型在输入噪声下的鲁棒性,揭示其脆弱性并提出改进方向。

Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation

  • 从输入扰动角度分类攻击与防御方法,系统评估模型表现。
  • 在噪声、混响等干扰下,语音转换质量显著下降,自然度和音色相似度降低。
  • 适合语音隐私保护、残障语音恢复等真实场景的研究者参考。

身份、口音、风格和情感是人类语音的重要组成部分。语音转换(VC)技术处理两个说话人的语音信号及提示词、情感标签等辅助信息,改变非语言特征而保留语义内容。近年来,VC模型在生成质量和个性化能力上快速进步,被广泛应用于隐私保护、逝者语音复现及构音障碍语音恢复。然而,由于训练数据干净,模型仅学习非鲁棒特征,导致在现实场景中面对噪声、混响、对抗攻击或微小扰动时性能显著下降。现有研究虽尝试探索攻击与防御策略,但对输入扰动下模型鲁棒性的全面理解仍不足。本文从输入操控视角分类现有攻击与防御方法,评估降质输入在可懂度、自然度、音色相似度和主观感知四个维度的影响,并指出开放问题与未来方向。

原文摘要 · Abstract (English)

Identity, accent, style, and emotions are essential components of human speech. Voice conversion (VC) techniques process the speech signals of two input speakers and other modalities of auxiliary information such as prompts and emotion tags. It changes para-linguistic features from one to another, while maintaining linguistic contents. Recently, VC models have made rapid advancements in both generation quality and personalization capabilities. These developments have attracted considerable attention for diverse applications, including privacy preservation, voice-print reproduction for the deceased, and dysarthric speech recovery. However, these models only learn non-robust features due to the clean training data. Subsequently, it results in unsatisfactory performances when dealing with degraded input speech in real-world scenarios, including additional noise, reverberation, adversarial attacks, or even minor perturbation. Hence, it demands robust deployments, especially in real-world settings. Although latest researches attempt to find potential attacks and countermeasures for VC systems, there remains a significant gap in the comprehensive understanding of how robust the VC model is under input manipulation. here also raises many questions: For instance, to what extent do different forms of input degradation attacks alter the expected output of VC models? Is there potential for optimizing these attack and defense strategies? To answer these questions, we classify existing attack and defense methods from the perspective of input manipulation and evaluate the impact of degraded input speech across four dimensions, including intelligibility, naturalness, timbre similarity, and subjective perception. Finally, we outline open issues and future directions.

语音转换鲁棒性语音安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。