arXiv:2512.18210cs.SDeess.SP2025-12ACL被引 4

通过优化数据多样性提升语音伪造检测泛化能力

A Data-Centric Approach to Generalizable Speech Deepfake Detection

  • 从数据多样性出发,设计可扩展的采样策略
  • 仅用3%数据即超越传统拼接方法,12k小时数据达顶尖性能
  • 适合需要高效、强泛化的语音安全研究者

语音深度伪造检测(SDD)的鲁棒泛化仍是主要挑战,模型常无法识别未见过的伪造方法。本文提出数据中心方法,从构建单一数据集和聚合多数据集两方面分析。首先开展大规模实证研究,量化源数据与生成器多样性的缩放规律;其次提出多样性优化采样策略(DOSS),包含剪枝版(DOSS-Select)与重加权版(DOSS-Weight)。实验表明,DOSS-Select仅使用总数据的3%即可超越朴素拼接基线;最终模型在12,000小时精选数据上采用最优DOSS-Weight策略,以更少数据和模型规模,在公开基准与新型商用API挑战集上均达到领先性能。

原文摘要 · Abstract (English)

Achieving robust generalization in speech deepfake detection (SDD) remains a primary challenge, as models often fail to detect unseen forgery methods. While research has focused on model-centric and algorithm-centric solutions, the impact of data composition is often underexplored. This paper proposes a data-centric approach, analyzing the SDD data landscape from two practical perspectives: constructing a single dataset and aggregating multiple datasets. To address the first perspective, we conduct a large-scale empirical study to characterize the data scaling laws for SDD, quantifying the impact of source and generator diversity. To address the second, we propose the Diversity-Optimized Sampling Strategy (DOSS), a principled framework for mixing heterogeneous data with two implementations: DOSS-Select (pruning) and DOSS-Weight (re-weighting). Our experiments show that DOSS-Select outperforms the naive aggregation baseline while using only 3% of the total available data. Furthermore, our final model, trained on a 12k-hour curated data pool using the optimal DOSS-Weight strategy, achieves state-of-the-art performance, outperforming large-scale baselines with greater data and model efficiency on both public benchmarks and a new challenge set of various commercial APIs.

语音伪造数据优化泛化能力检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。