捕捉音频伪造中多特征协同模式,提升检测精度。
HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
- 用超图建模多个特征的高阶关联关系。
- 平均错误率降低12.68%,部分场景达22.15%。
- 适合对抗复杂伪造的检测任务。
AIGC技术的进步使得合成高度逼真的音频深度伪造成为可能,能骗过人类听觉感知。尽管已有众多音频深度伪造检测(ADD)方法,但多数依赖局部时序或频谱特征,或成对关系,忽视了高阶交互(HOIs)。HOIs能捕捉超出单个特征贡献的判别性模式。我们提出HyperPotter,一种基于超图的框架,通过聚类生成的超边与类别感知原型初始化,捕获协同模式相关的高阶关系。在13个测试集上的大量实验表明,HyperPotter在11个数据集上优于基线,所有测试集平均相对EER降低12.68%,改善集上达22.15%。结果表明其具备强大的跨场景泛化能力,但在严重编码器或信道失真下仍存在鲁棒性局限。
原文摘要 · Abstract (English)
Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception. Although numerous audio deepfake detection (ADD) methods have been developed, most rely on local temporal/spectral features or pairwise relations, overlooking high-order interactions (HOIs). HOIs capture discriminative patterns that emerge from multiple feature components beyond their individual contributions. We propose HyperPotter, a hypergraph-based framework designed to capture high-order relations associated with synergistic patterns through clustering-based hyperedges with class-aware prototype initialization. Extensive experiments on 13 test sets show that HyperPotter improves over the baseline on 11 sets, yielding an average relative EER reduction of 12.68\% across all test sets and 22.15\% on the improved sets. These results demonstrate strong cross-scenario generalization, while also revealing robustness limits under severe codec or channel distortion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。