arXiv:2509.13085eess.AS2025-09中稿 · IEEE ASRU 2025

用可学习的声学吸引子提升语音欺骗定位与分类精度

Token-based Attractors and Cross-attention in Spoof Diarization

  • 引入可学习令牌表征真实与伪造语音特征
  • 在PartialSpoof数据集上优于现有方法
  • 适合语音安全与反欺骗研究者参考

语音欺骗分段旨在识别给定语音中「何时被伪造」,通过定位伪造区域并确定其篡改技术。作为该任务的初步探索,已有工作提出双分支模型实现定位与欺骗类型聚类,为语音欺骗分段奠定基础。然而其结构简单,难以捕捉复杂欺骗模式,且缺乏明确参考点以区分真实语音与各类伪造语音。为此,本文提出基于可学习令牌的方法,每个令牌代表真实或伪造语音的声学特征,通过与帧级嵌入交互提取判别性表示,增强真实与生成语音的分离能力。在PartialSpoof数据集上的大量实验表明,该方法在真实语音检测与欺骗方式聚类任务上均优于现有方法。

原文摘要 · Abstract (English)

Spoof diarization identifies ``what spoofed when" in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for localization and spoof type clustering, which laid the foundation for spoof diarization. However, its simple structure limits the ability to capture complex spoofing patterns and lacks explicit reference points for distinguishing between bona fide and various spoofing types. To address these limitations, our approach introduces learnable tokens where each token represents acoustic features of bona fide and spoofed speech. These attractors interact with frame-level embeddings to extract discriminative representations, improving separation between genuine and generated speech. Vast experiments on PartialSpoof dataset consistently demonstrate that our approach outperforms existing methods in bona fide detection and spoofing method clustering.

语音安全欺骗检测注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。