arXiv:2509.14959eess.AScs.AI2025-09

用最优传输对齐语音分布,实现无需模型参数的音频欺骗攻击

Discrete optimal transport is a strong audio adversarial attack

  • 通过熵正则最优传输与top-k投影,将生成语音嵌入对齐真实语音分布
  • 在ASVspoof2019和ASVspoof5上显著提升反欺骗系统误报率,降低识别性能
  • 攻击不依赖梯度或训练数据,可跨数据集迁移且对抗微调仍有效

本文研究离散最优传输(DOT)作为黑盒攻击方法,针对现代自动说话人验证(ASV)和防欺骗检测(CM)系统。该攻击作为后处理分布对齐步骤:将生成语音(或他人语音)的帧级WavLM嵌入,通过熵正则最优传输与top-k重心投影,对齐到无配对的真实语音池,随后进行神经声码合成。与基于梯度的攻击不同,该方法无需访问模型参数、梯度或训练数据。在ASVspoof2019和ASVspoof5上的实验表明,DOT攻击显著提高CM的等错误率(EER),并大幅削弱多种欺骗攻击下的ASV性能。攻击具备跨数据集迁移能力,且在CM微调后仍有效。基于说话人相似性、Fréchet音频距离及嵌入分布可视化分析表明,DOT成功的关键在于将源语音推向表示空间中的真实语音区域,而非最大化说话人相似性。结果表明,基于最优传输的分布对齐是当前ASV与反欺骗系统中尚未被充分探索的攻击路径。

原文摘要 · Abstract (English)

In this paper, we investigate discrete optimal transport (DOT) as a black-box attack against modern automatic speaker verification (ASV) and anti-spoofing countermeasure (CM) systems. Our attack operates as a post-processing distribution-alignment step. Frame-level WavLM embeddings of generated speech (or another person speech) are aligned to an unpaired bona fide speech pool using entropic optimal transport and a top-k barycentric projection, followed by neural vocoding. Unlike gradient-based attacks, the proposed method requires no access to model parameters, gradients, or training data. Experiments on ASVspoof2019 and ASVspoof5 demonstrate that DOT attack substantially increases CM EER and substantially degrades ASV performance across multiple spoofing attacks. The attack transfers across datasets and remains effective after CM fine-tuning. Analysis using speaker similarity, Fréchet Audio Distance, and visualization of embedding distributions suggests that DOT succeeds by shifting source speech toward bona fide regions of the representation space rather than by maximizing speaker similarity. These results indicate that optimal-transport-based distribution alignment represents a previously underexplored attack vector for contemporary ASV and anti-spoofing systems.

音频安全对抗攻击最优传输说话人验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。