arXiv:2603.02781cs.CRcs.AI2026-03被引 2

通过特征对齐提升语音伪造攻击效率,仅用50次查询即可成功。

Scores Know Bobs Voice: Speaker Impersonation Attack

  • 构建逆映射模型,使生成潜空间与语音识别特征空间对齐。
  • 实现平均10倍查询减少,50次查询下成功率最高达91.65%。
  • 支持新型投影攻击,适用于评估现代语音识别系统的鲁棒性。

深度学习推动了说话人识别系统(SRS)的广泛应用,但其仍易受基于得分的伪造攻击。现有在原始波形上操作的攻击需大量查询,因高维音频空间优化困难。生成模型中的潜在空间优化虽更高效,但其结构由数据分布决定,不天然包含说话人判别几何特征,导致优化轨迹难以对准提升目标得分的对抗方向。为此,我们提出一种基于反演的生成式攻击框架,显式将合成模型的潜空间与SRS的判别特征空间对齐。我们分析了反演模型在得分攻击中的需求,并引入特征对齐的反演策略,实现潜表示与说话人嵌入的几何同步。该对齐确保潜空间更新直接转化为得分提升,并开启新攻击范式,如子空间投影攻击,此前因缺乏精确特征到音频映射而不可行。实验表明,本方法显著提升查询效率,平均仅需前序方法1/10的查询量即达竞争性攻击成功率。特别地,所启用的子空间投影攻击在仅50次查询下达到91.65%成功率。这些结果确立特征对齐反演是评估现代SRS对得分伪造威胁鲁棒性的关键工具。

原文摘要 · Abstract (English)

Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing attacks that operate directly on raw waveforms require a large number of queries due to the difficulty of optimizing in high-dimensional audio spaces. Latent-space optimization within generative models offers improved efficiency, but these latent spaces are shaped by data distribution matching and do not inherently capture speaker-discriminative geometry. As a result, optimization trajectories often fail to align with the adversarial direction needed to maximize victim scores. To address this limitation, we propose an inversion-based generative attack framework that explicitly aligns the latent space of the synthesis model with the discriminative feature space of SRSs. We first analyze the requirements of an inverse model for score-based attacks and introduce a feature-aligned inversion strategy that geometrically synchronizes latent representations with speaker embeddings. This alignment ensures that latent updates directly translate into score improvements. Moreover, it enables new attack paradigms, including subspace-projection-based attacks, which were previously infeasible due to the absence of a faithful feature-to-audio mapping. Experiments show that our method significantly improves query efficiency, achieving competitive attack success rates with on average 10x fewer queries than prior approaches. In particular, the enabled subspace-projection-based attack attains up to 91.65% success using only 50 queries. These findings establish feature-aligned inversion as a key tool for evaluating the robustness of modern SRSs against score-based impersonation threats.

语音伪造对抗攻击生成模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。