arXiv:2608.09593cs.SDcs.AI2026-08

首个区分语音与环境音的音频伪造检测基准,揭示两类音频检测差异。

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

  • 将语音与环境音分开展示,实现组件级检测评估
  • 环境音伪造比语音更易被检测,现有模型在两者上均失效
  • 适用于需区分音频成分的反伪造研究者

近年来语音合成与音频生成技术进步使高保真声学伪造成本降低且难以溯源,形成语音与背景音分别篡改却视频真实的真实攻击场景。现有研究或聚焦视觉篡改,或孤立处理语音检测,或将语音与非语音音频合并为单一音频流,忽视了背景音带来的独特取证挑战。这种混淆具有严重后果:两类声音源于不同生成机制,呈现不同伪影特征,对检测系统构成不同难题。本文提出MADBench,首个将语音与环境音视为独立声学成分的基准,支持对独立篡改来源的组件感知检测评估。我们采用统一协议测试主流检测器与多模态大语言模型。实验发现,通用编码器下环境音伪造比合成语音更易检测;现有预训练检测器在两类成分上均表现不佳;且篡改环境音会不对称地削弱语音伪造检测效果——这些现象在以往单标签基准中完全不可见。MADBench为未来鲁棒、组件感知的音频伪造检测研究奠定了严谨基础。

原文摘要 · Abstract (English)

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.

音频伪造检测基准多模态深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。