arXiv:2606.05101cs.SDcs.LG2026-06中稿 · ICML被引 1

用大模型自动生成骗过语音深伪检测器的音频,无需人工干预。

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

论文配图:FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors
图 1 · 摘自论文原文
  • 利用大模型的上下文学习能力,自动探索语音生成空间。
  • 生成样本使检测器误判率最高提升94%,且攻击可跨模型迁移。
  • 适合安全研究者和检测模型开发者快速评估系统漏洞。

语音深伪检测(ADD)模型对防范文本转语音(TTS)模型的恶意使用至关重要。评估与强化ADD模型需构建覆盖生成音频空间并突出高错误区域的数据集。现有数据集构建方法面临两大挑战:(i) 依赖人工收集,(ii) 难以高效发现ADD模型的盲区。为此,我们提出FoeGlass,首个针对ADD的黑盒自动化红队方法,能有效发现当前主流深伪基准未充分覆盖的检测失败模式。FoeGlass利用大模型的上下文学习能力,在仅拥有所有组件黑盒访问权限的前提下,探索TTS模型输入空间,生成可欺骗目标ADD的音频样本。通过精心设计基于多样性度量的上下文,缓解了自动化红队系统常见的模式坍缩问题。在多个开源ADD与TTS模型上的实证评估表明,FoeGlass生成的数据显著降低误报率,相比无条件采样基线和近期伪造数据集最高提升94%;且无需人工监督。此外,FoeGlass生成的攻击具备跨目标ADD的可迁移性,体现其广泛适用性与易用性。最后,将ADD模型在FoeGlass生成样本上微调,使其鲁棒性提升达41%。

原文摘要 · Abstract (English)

Audio deepfake detection (ADD) models are critical for countering the malicious use of text-to-speech (TTS) models. Evaluating and strengthening ADD models requires developing datasets that span the space of generated audio and highlight high-error regions. Existing dataset development strategies face two challenges: (i) manual collection, and (ii) inefficient discovery of blind spots in the ADD models. To address these challenges, we propose FoeGlass, the first black-box automated red-teaming method for ADDs, which effectively discovers ADD failure modes in the space of generated audio underexplored by state-of-the-art deepfake benchmarks. FoeGlass uses the in-context learning capabilities of an LLM to explore the input space of a TTS model, generating audio samples that fool the target ADD using only black-box access to all components. By using a carefully designed context based on diversity measurements, FoeGlass mitigates the common problem of mode collapse in automated red-teaming systems. Empirical evaluations on several open-source ADD and TTS models demonstrate that data generated from FoeGlass substantially improves the false negative rates over unconditional sampling baselines and recent spoofing datasets by up to 94%, while requiring no manual supervision. Furthermore, we show that the attacks generated by FoeGlass are transferable across different target ADDs, demonstrating its broad applicability and ease of use for the automated red teaming of ADD systems. Finally, fine-tuning ADD models on FoeGlass-generated samples notably enhances the robustness of the detectors (up 41%).

语音生成深度伪造红队测试大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。