arXiv:2605.00251cs.SDcs.CL2026-05中稿 · ICML被引 1

提出首个语音伪造检测通用编码器Alethia,显著提升抗干扰与跨域泛化能力。

Alethia: A Foundational Encoder for Voice Deepfakes

论文配图:Alethia: A Foundational Encoder for Voice Deepfakes
图 1 · 摘自论文原文
  • 结合瓶颈嵌入预测与流匹配谱图重建进行预训练
  • 在56个数据集上超越现有模型,零样本泛化至歌唱伪造
  • 揭示连续嵌入预测优于离散目标,适合伪造检测研究者

现有语音伪造检测与定位模型严重依赖语音基础模型(SFM)提取的表征,但微调已进入收益递减阶段。本文转向预训练,提出一种新方法:将瓶颈嵌入预测与基于流匹配的谱图重建相结合。由此得到的Alethia是首个面向多种语音伪造检测与定位任务的基础音频编码器。我们在5个不同任务、56个基准数据集上评估,结果表明Alethia显著优于当前最优的SFM,在真实世界扰动下表现更鲁棒,并能实现对未见领域(如歌唱伪造)的零样本泛化。我们还发现离散目标在掩码标记预测中存在局限,强调连续嵌入预测与生成式预训练对捕捉伪造痕迹的重要性。

原文摘要 · Abstract (English)

Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on $5$ different tasks with $56$ benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.

语音伪造基础模型零样本音频编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。