arXiv:2602.16343cs.SDcs.LG2026-02中稿 · ICASSP 2026

研究神经音频编码器在语音伪造检测中的双重角色,提出新标注策略。

How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection

  • 用神经音频编码器生成语音伪造数据,模拟攻击者行为
  • 发现不同标注方式导致检测性能差异显著
  • 适用于语音安全、伪造检测研究者

由于文本转语音系统通常不直接生成波形,近期的欺骗检测研究使用声码器和神经音频编码器重合成波形来模拟攻击者。与专为语音合成设计的声码器不同,神经音频编码器最初用于音频压缩存储与传输。其语音离散化能力也激发了基于语言模型的语音合成研究。由于具备双重功能,编码器重合成的数据可能被标记为真实或伪造。目前对此问题的研究极少。本研究为此构建了更具挑战性的ASVspoof 5数据集扩展版本,探讨不同标注方式对检测性能的影响,并提供标注策略的实用建议。

原文摘要 · Abstract (English)

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically designed for speech synthesis, neural audio codecs were originally developed for compressing audio for storage and transmission. However, their ability to discretize speech also sparked interest in language-modeling-based speech synthesis. Owing to this dual functionality, codec resynthesized data may be labeled as either bonafide or spoof. So far, very little research has addressed this issue. In this study, we present a challenging extension of the ASVspoof 5 dataset constructed for this purpose. We examine how different labeling choices affect detection performance and provide insights into labeling strategies.

语音伪造检测编码器数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。