arXiv:2603.14033cs.SDcs.AI2026-03

语音修复与音质转换让假声检测失效,需区分原始真伪与处理状态。

What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection

  • 将音频真伪标签拆分为源真实性和处理状态,更准确刻画音频变化
  • 模型能可靠识别是否被处理,但难以区分处理后的真声与假声
  • 适用于需要细粒度音频安全评估的场景,如司法鉴定、语音助手

语音反欺骗系统通常对整个语音片段赋予单一真实性标签。当变换不改变说话人身份和语言内容时,这一设定变得模糊。本文研究了良性、保真性不变的语音变换,包括音质转换和语音修复,应用于真实与伪造语音。我们不再将所有处理后音频视为伪造,而是将标签分解为源真实性与处理状态。在自监督学习(SSL)表示和DF-Arena微调实验中发现,处理状态比源归属具有更强的可迁移性:检测器往往能识别语音已被处理,但仍混淆处理后的真实语音与处理后的伪造语音。结果表明,语音深度伪造防御必须超越二元真假范式。鲁棒检测需提供关于源真实性、处理状态及处理位置的细粒度报告。

原文摘要 · Abstract (English)

Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance. This formulation becomes under-specified for transformations where the underlying speaker identity and linguistic content remain unchanged. We study this problem using benign, authenticity-preserving speech transformations, including voice quality conversion and speech restoration, applied to both bona fide and spoofed speech. Instead of treating all processed audio as spoofed, we factorise labels into source authenticity and processed status. Across SSL representations and DF-Arena fine-tuning experiments, we find that utterance processing status can transfer more reliably than source attribution: detectors can often identify that speech has been processed, while still confusing processed bona fide and processed spoofed speech. These results suggest that audio deepfake defences must move beyond the binary spoofed/authentic paradigm. Robust detection requires granular reporting on source authenticity, processing status, and precise processing localisation.

语音伪造深度伪造音频安全细粒度检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。