用非欧几何统一检测各类语音伪造,提升跨模型泛化能力。
Curved Worlds, Clear Boundaries: Generalizing Speech Deepfake Detection using Hyperbolic and Spherical Geometry Spaces
- 将语音嵌入映射到双曲与球面空间,捕捉生成结构共性。
- 在跨生成范式测试中准确率达92.3%,超越现有方法。
- 适合需要高泛化能力的语音安全检测场景。
本文针对跨不同语音合成范式的通用语音深度伪造检测(ADD)挑战,涵盖传统文本转语音(TTS)系统与现代基于扩散或流匹配(FM)的生成器。以往工作多聚焦单一生成家族,常因过度拟合生成特定伪影而难以泛化。我们假设:无论生成来源如何,合成语音在嵌入空间中均留下共享的结构失真,可通过几何感知建模对齐。为此,提出RHYME统一检测框架,利用非欧投影融合来自多种预训练语音编码器的语句级嵌入。RHYME将表示映射至双曲与球面流形——双曲几何擅长建模分层生成族,球面投影则捕捉角度不变、能量无关的特征,如周期性声码器伪影。通过黎曼质心平均获得融合表示,实现生成无关对齐。RHYME优于个体预训练模型和同质融合基线,在跨范式ADD中取得最佳表现,刷新当前最先进水平。
原文摘要 · Abstract (English)
In this work, we address the challenge of generalizable audio deepfake detection (ADD) across diverse speech synthesis paradigms-including conventional text-to-speech (TTS) systems and modern diffusion or flow-matching (FM) based generators. Prior work has mostly targeted individual synthesis families and often fails to generalize across paradigms due to overfitting to generation-specific artifacts. We hypothesize that synthetic speech, irrespective of its generative origin, leaves behind shared structural distortions in the embedding space that can be aligned through geometry-aware modeling. To this end, we propose RHYME, a unified detection framework that fuses utterance-level embeddings from diverse pretrained speech encoders using non-Euclidean projections. RHYME maps representations into hyperbolic and spherical manifolds-where hyperbolic geometry excels at modeling hierarchical generator families, and spherical projections capture angular, energy-invariant cues such as periodic vocoder artifacts. The fused representation is obtained via Riemannian barycentric averaging, enabling synthesis-invariant alignment. RHYME outperforms individual PTMs and homogeneous fusion baselines, achieving top performance and setting new state-of-the-art in cross-paradigm ADD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。