arXiv:2604.17642eess.AS2026-04ACL被引 1

针对病理性语音伪造检测,提出新基准与高效识别框架。

HCFD: A Benchmark for Audio Deepfake Detection in Healthcare

论文配图:HCFD: A Benchmark for Audio Deepfake Detection in Healthcare
图 1 · 摘自论文原文
  • 构建病理语音伪造数据集,聚焦编码器生成的合成语音
  • 提出几何感知模型PHOENIX-Mamba,在多病种上达97%准确率
  • 适合医疗语音安全、深度伪造检测研究者参考

本文提出医疗编码器伪造检测(HCFD)新任务,聚焦病理语音条件下的编码器合成语音检测。首先发布首个面向病理场景的语音伪造数据集Healthcare CodecFake,包含多种临床病症及编码器类型下的真实与非病理合成语音对。实验表明,主流健康语音训练的伪造检测模型在该数据集上表现不佳,凸显专用模型需求。其次,发现PaSST在该任务中优于现有模型,因其基于分块的时频表示。最后,提出PHOENIX-Mamba框架,将伪造模式建模为双曲空间中的自发现多模式,实现跨病症与编码器的最优性能:在E-Dep、E-Alz、E-Dys上分别达到97.04%、96.73%、96.57%准确率;中文数据集上也保持高精度(94.41、94.40、93.20)。该几何建模方式增强了对病理语音变异的鲁棒性。

原文摘要 · Abstract (English)

In this study, we present Healthcare Codec-Fake Detection (HCFD), a new task for detecting codec-fakes under pathological speech conditions. We intentionally focus on codec based synthetic speech in this work, since neural codec decoding forms a core building block in modern speech generation pipelines. First, we release Healthcare CodecFake, the first pathology-aware dataset containing paired real and NAC-synthesized speech across multipl clinical conditions and codec families. Our evaluations show that SOTA codec-fake detectors trained primarily on healthy speech perform poorly on Healthcare CodecFake, highlighting the need for HCFD-specific models. Second, we demonstrate that PaSST outperforms existing speech-based models for HCFD, benefiting from its patch-based spectro-temporal representation. Finally, we propose PHOENIX-Mamba, a geometry-aware framework that models codec-fakes as multiple self-discovered modes in hyperbolic space and achieves the strongest performance on HCFD across clinical conditions and codecs. Experiments on HCFK show that PHOENIX-Mamba (PaSST) achieves the best overall performance, reaching 97.04 Acc on E-Dep, 96.73 on E-Alz, and 96.57 on E-Dys, while maintaining strong results on Chinese with 94.41 (Dep), 94.40 (Alz), and 93.20 (Dys). This geometry-aware formulation enables self-discovered clustering of heterogeneous codec-fake modes in hyperbolic space, facilitating robust discrimination under pathological speech variability. PHOENIX-Mamba achieves topmost performance on the HCFD task across clinical conditions and codecs.

语音伪造医疗AI双曲空间检测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。