arXiv:2508.01897cs.SDeess.AS2025-08中稿 · publication on Int…被引 6

用庞加莱球构建层级特征,提升语音伪造检测泛化能力

Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere

  • 在庞加莱球中学习攻击类型的分层结构,捕捉隐藏的语义关系
  • 在四个数据集上等错误率优于现有方法,最高提升12.3%
  • 适合需要跨域鲁棒性的语音安全研究者

语音深度伪造检测面临真实世界攻击多样性和域差异带来的严重泛化挑战。现有方法多依赖欧氏距离,难以有效捕捉攻击类别与域因素相关的内在层级结构。为此,我们提出Poin-HierNet框架,在庞加莱球中构建域不变的层次化表示。该框架包含三个核心组件:1)庞加莱原型学习(PPL),通过多个数据原型对齐样本特征,捕获超越人工标签的多层次结构;2)层次结构学习(HSL),利用顶层原型建立树状层次结构;3)庞加莱特征去相关(PFW),通过特征去相关抑制域敏感特征以增强域不变性。我们在ASVspoof 2019 LA、ASVspoof 2021 LA、ASVspoof 2021 DF和In-The-Wild四个数据集上评估,实验结果表明,Poin-HierNet在等错误率(EER)上显著优于当前最优方法。

原文摘要 · Abstract (English)

Audio deepfake detection (ADD) faces critical generalization challenges due to diverse real-world spoofing attacks and domain variations. However, existing methods primarily rely on Euclidean distances, failing to adequately capture the intrinsic hierarchical structures associated with attack categories and domain factors. To address these issues, we design a novel framework Poin-HierNet to construct domain-invariant hierarchical representations in the Poincaré sphere. Poin-HierNet includes three key components: 1) Poincaré Prototype Learning (PPL) with several data prototypes aligning sample features and capturing multilevel hierarchies beyond human labels; 2) Hierarchical Structure Learning (HSL) leverages top prototypes to establish a tree-like hierarchical structure from data prototypes; and 3) Poincaré Feature Whitening (PFW) enhances domain invariance by applying feature whitening to suppress domain-sensitive features. We evaluate our approach on four datasets: ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2021 DF, and In-The-Wild. Experimental results demonstrate that Poin-HierNet exceeds state-of-the-art methods in Equal Error Rate.

语音伪造深度伪造庞加莱空间域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。