arXiv:2604.01083cs.SDcs.AI2026-04被引 6

无需训练即可检测部分音频伪造,通过分析语音模型嵌入轨迹的动态变化。

TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models

  • 利用冻结语音基础模型的嵌入轨迹变化,捕捉伪造拼接处的突变特征。
  • 在多个基准上达到8.08%的等错误率,优于部分监督基线。
  • 无需标注数据或重新训练,适合快速应对新型生成模型的伪造攻击。

部分音频伪造(合成片段拼接至真实录音)极具欺骗性,因多数内容仍为真实。现有检测方法依赖帧级标注,易过拟合特定生成管道,且需随新模型重训。本文提出无需训练的TRACE框架,基于语音基础模型隐含的取证信号:真实语音嵌入轨迹平滑连续,而拼接边界会导致帧间过渡突变。通过分析冻结模型表示的一阶动态,实现无训练、无标注、无结构修改的检测。在涵盖两种语言的四个基准上评估,使用六种语音基础模型。在PartialSpoof上达到8.08% EER,媲美微调监督基线;在最具挑战性的LlamaPartialSpoof(LLM驱动商业合成)上,以24.12%对比24.49%的EER超越监督基线,且未使用目标域数据。结果表明,语音基础模型的时间动态可作为泛化有效的无训练取证信号。

原文摘要 · Abstract (English)

Partial audio deepfakes, where synthesized segments are spliced into genuine recordings, are particularly deceptive because most of the audio remains authentic. Existing detectors are supervised: they require frame-level annotations, overfit to specific synthesis pipelines, and must be retrained as new generative models emerge. We argue that this supervision is unnecessary. We hypothesize that speech foundation models implicitly encode a forensic signal: genuine speech forms smooth, slowly varying embedding trajectories, while splice boundaries introduce abrupt disruptions in frame-level transitions. Building on this, we propose TRACE (Training-free Representation-based Audio Countermeasure via Embedding dynamics), a training-free framework that detects partial audio deepfakes by analyzing the first-order dynamics of frozen speech foundation model representations without any training, labeled data, or architectural modification. We evaluate TRACE on four benchmarks that span two languages using six speech foundation models. In PartialSpoof, TRACE achieves 8.08% EER, competitive with fine-tuned supervised baselines. In LlamaPartialSpoof, the most challenging benchmark featuring LLM-driven commercial synthesis, TRACE surpasses a supervised baseline outright (24.12% vs. 24.49% EER) without any target-domain data. These results show that temporal dynamics in speech foundation models provide an effective, generalize signal for training-free audio forensics.

音频伪造无训练检测语音模型嵌入轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。