构建多语言高保真假语音检测数据集,解决真实感缺失问题
HQ-MPSD: A Multilingual Artifact-Controlled Benchmark for Partial Deepfake Speech Detection
- 基于细粒度强制对齐选取语义连贯拼接点,减少听觉可见伪影
- 涵盖8语言550人共350.8小时语音,跨语言检测性能下降超80%
- 适合研究真实场景下深度伪造语音的鲁棒检测方法
检测局部深度伪造语音极具挑战性,因篡改仅发生在短片段而周围音频保持真实。现有检测方法受限于数据集质量,许多数据依赖过时合成系统,引入非真实的人工痕迹而非自然篡改线索。为此,我们提出HQ-MPSD——一个高质量多语言局部深度伪造语音数据集。该数据集通过细粒度强制对齐获取语言连贯的拼接点,保留韵律与语义连续性,最小化听觉与视觉边界伪影。数据集包含8种语言、550名说话人,总计350.8小时语音,并加入背景音效以更贴近真实声学环境。主观评分(MOS)与频谱分析验证样本具有高感知自然度。我们在跨语言与跨数据集条件下基准测试先进检测模型,所有模型在HQ-MPSD上的性能下降均超过80%。结果表明,去除低层伪影并引入多语言与声学多样性后,模型泛化能力面临严峻挑战,为局部深度伪造语音检测提供了更真实、更具难度的评估基准。数据集地址:https://zenodo.org/records/17929533。
原文摘要 · Abstract (English)
Detecting partial deepfake speech is challenging because manipulations occur only in short regions while the surrounding audio remains authentic. However, existing detection methods are fundamentally limited by the quality of available datasets, many of which rely on outdated synthesis systems and generation procedures that introduce dataset-specific artifacts rather than realistic manipulation cues. To address this gap, we introduce HQ-MPSD, a high-quality multilingual partial deepfake speech dataset. HQ-MPSD is constructed using linguistically coherent splice points derived from fine-grained forced alignment, preserving prosodic and semantic continuity and minimizing audible and visual boundary artifacts. The dataset contains 350.8 hours of speech across eight languages and 550 speakers, with background effects added to better reflect real-world acoustic conditions. MOS evaluations and spectrogram analysis confirm the high perceptual naturalness of the samples. We benchmark state-of-the-art detection models through cross-language and cross-dataset evaluations, and all models experience performance drops exceeding 80% on HQ-MPSD. These results demonstrate that HQ-MPSD exposes significant generalization challenges once low-level artifacts are removed and multilingual and acoustic diversity are introduced, providing a more realistic and demanding benchmark for partial deepfake detection. The dataset can be found at: https://zenodo.org/records/17929533.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。