arXiv:2608.13817eess.AScs.LG2026-08

利用语音生理约束的时序动态,提升自监督模型对音频伪造的检测能力。

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

论文配图:Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection
图 1 · 摘自论文原文
  • 通过因果LSTM建模自监督特征的时序轨迹,捕捉真实语音的生理约束
  • 在跨语料库测试中,动态方法显著优于静态方法,最高误报率降至0.75%
  • 仅用真实语音训练即可超越已有监督基线,适合对抗多样伪造技术

人类语音生成受生理限制,导致声学信号具有特定时序结构。我们假设这些约束会在自监督学习(SSL)模型的隐空间中体现为有结构的轨迹动态,而合成语音会破坏这种规律,可被检测。为此,我们仅使用真实语音训练因果LSTM帧预测器(阶段1),采用专用于反伪造的Wav2Vec2-Large-AntiDeepfake骨干网络,并与相同特征的全局平均池化基线对比,以分离时序建模的贡献。阶段2引入监督多层感知机,冻结LSTM内部状态进行分类,以分析欺骗标注的作用。系统在六个基准上表现优异:ASVspoof 2019/2021、Codecfake、In-the-Wild、MLAAD-EN和Deepfake-Eval-2024,其中在ASVspoof 2021上达到0.75%的最佳公开误报率;值得注意的是,仅用真实语音训练的阶段1已超越同骨干网络的已有监督基线,在DE2024上达30.35%。近域任务中静态与动态方法表现相近,但在更难的跨语料库任务(涵盖多种合成方法)中,轨迹动态带来显著提升,证实了时序生理约束蕴含超越语句级统计的检测信号。

原文摘要 · Abstract (English)

Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals. We hypothesise that these constraints manifest as structured trajectory dynamics in the latent space of Self-Supervised Learning (SSL) models, and that synthetic speech violates them detectably. To test this hypothesis, we train a causal Long Short-Term Memory (LSTM) next-frame predictor on bonafide speech only (Stage 1), using the deepfake-specialised SSL backbone Wav2Vec2-Large-AntiDeepfake, and compare against a static global-average-pooling baseline using identical features, thus isolating the contribution of temporal modelling. A supervised Stage 2, which trains a Multi-Layer Perceptron on the frozen LSTM internal states using labelled data, is included to characterise the role of spoof supervision. Our system achieves competitive or state-of-the-art performance across six benchmarks: ASVspoof 2019/2021, Codecfake, In-the-Wild, MLAAD-EN, and Deepfake-Eval-2024, including best published EER on ASVspoof 2021 (0.75\%) and, notably, Stage 1 trained on bonafide speech only surpasses the published supervised baseline from the same backbone on DE2024 (30.35\%). On near-domain benchmarks, static and dynamic approaches perform comparably. On harder cross-corpus benchmarks with diverse synthesis methods, trajectory dynamics provide substantial gains, confirming that temporal physiological constraints carry detection signal beyond utterance-level statistics.

音频伪造自监督学习时序建模检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。