arXiv:2506.22783cs.CVcs.AI2025-06被引 5

用语言推理精准篡改语音片段,让假音频更逼真且易被检测。

PhonemeFake: Redefining Deepfake Realism with Language-Driven Segmental Manipulation and Adaptive Bilevel Detection

  • 通过语言理解定位关键语音段进行篡改,提升伪造真实感。
  • 使人类识别错误率最高降低42%,检测模型误报率下降91%。
  • 开源数据集与自适应检测模型,适合安全研究与防御开发。

深度伪造攻击随生成模型进步日益严峻。但本研究发现,现有深度伪造数据集无法有效欺骗人类感知,与真实攻击影响公众舆论的威胁不符,凸显构建更逼真攻击向量的必要性。我们提出PhonemeFake(PF)攻击方法,利用语言推理操控关键语音片段,使人类感知准确率下降最高达42%,基准检测模型准确率下降最高达94%。我们已在HuggingFace发布可复现的PF数据集,并开源自适应双层检测模型,该模型能动态分配计算资源至被篡改区域。在三个已知深度伪造数据集上的广泛实验表明,该检测模型将等错误率(EER)降低91%,同时实现最高90%的速度提升,计算开销极低,定位精度优于现有方法,具备可扩展性。

原文摘要 · Abstract (English)

Deepfake (DF) attacks pose a growing threat as generative models become increasingly advanced. However, our study reveals that existing DF datasets fail to deceive human perception, unlike real DF attacks that influence public discourse. It highlights the need for more realistic DF attack vectors. We introduce PhonemeFake (PF), a DF attack that manipulates critical speech segments using language reasoning, significantly reducing human perception by up to 42% and benchmark accuracies by up to 94%. We release an easy-to-use PF dataset on HuggingFace and open-source bilevel DF segment detection model that adaptively prioritizes compute on manipulated regions. Our extensive experiments across three known DF datasets reveal that our detection model reduces EER by 91% while achieving up to 90% speed-up, with minimal compute overhead and precise localization beyond existing models as a scalable solution.

深度伪造语音伪造自适应检测语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。