针对环境音与语音分量伪造难题,提出新数据集与检测挑战。
ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Evaluation Plan
- 构建分量级音频反伪造数据集CompSpoofV2,含250万样本
- 提出分离增强联合学习框架,提升对混合音的检测能力
- 面向真实场景的语音与环境音协同伪造检测,适合安全与内容可信研究者
真实环境中录制的音频通常包含前景语音与背景环境声的混合。随着文本转语音、语音转换等生成模型的快速发展,可独立修改其中任一成分。此类分量级篡改更难检测,因未被篡改的成分可能误导原有全音频深度伪造检测系统,且对人类听觉更自然。为填补此空白,我们提出CompSpoofV2数据集及分离增强的联合学习框架。CompSpoofV2是大规模精心标注的数据集,包含超过250,000个音频样本,总时长约283小时。基于该数据集与框架,我们发起环境感知语音与声音深度伪造检测挑战(ESDD2),聚焦分量级伪造,即语音与环境声均可被篡改或合成,营造更具挑战性与真实性的检测场景。该挑战将与IEEE多媒体与博览会2026(ICME 2026)同步举行。
原文摘要 · Abstract (English)
Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be modified independently. Such component-level manipulations are harder to detect, as the remaining unaltered component can mislead the systems designed for whole deepfake audio, and they often sound more natural to human listeners. To address this gap, we have proposed CompSpoofV2 dataset and a separation-enhanced joint learning framework. CompSpoofV2 is a large-scale curated dataset designed for component-level audio anti-spoofing, which contains over 250k audio samples, with a total duration of approximately 283 hours. Based on the CompSpoofV2 and the separation-enhanced joint learning framework, we launch the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2), focusing on component-level spoofing, where both speech and environmental sounds may be manipulated or synthesized, creating a more challenging and realistic detection scenario. The challenge will be held in conjunction with the IEEE International Conference on Multimedia and Expo 2026 (ICME 2026).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。