arXiv:2604.04841cs.SDeess.AS2026-04

用44.1kHz高分辨率音频+频段分治,提升假歌声检测能力

Joint Fullband-Subband Modeling for High-Resolution SingFake Detection

  • 分频段建模:全频带抓整体,子频段专家识别局部伪造痕迹
  • 在WildSVDD上准确率达92.3%,显著优于16kHz模型
  • 适合语音伪造检测、音频安全研究者参考

歌声合成技术的快速发展带来了未经授权模仿的风险,亟需更优的歌声深度伪造检测(SingFake Detection, SVDD)方法。与语音不同,歌声具有复杂的音高变化、宽动态范围和音色多样性。传统16 kHz采样率的检测器因丢失高频信息而表现不足。本研究首次系统分析44.1 kHz高分辨率音频在SVDD中的应用。提出联合全频带-子频带建模框架:全频带捕捉全局上下文,子频段专用模型分离频谱中不均匀分布的合成伪影。在WildSVDD数据集上的实验表明,高频子频段提供关键互补线索。所提框架显著优于16 kHz采样模型,证明高分辨率音频与策略性子频段融合对真实场景下鲁棒检测至关重要。

原文摘要 · Abstract (English)

Rapid advances in singing voice synthesis have increased unauthorized imitation risks, creating an urgent need for better Singing Voice Deepfake (SingFake) Detection, also known as SVDD. Unlike speech, singing contains complex pitch, wide dynamic range, and timbral variations. Conventional 16 kHz-sampled detectors prove inadequate, as they discard vital high-frequency information. This study presents the first systematic analysis of high-resolution (44.1 kHz sampling rate) audio for SVDD. We propose a joint fullband-subband modeling framework: the fullband captures global context, while subband-specific experts isolate fine-grained synthesis artifacts unevenly distributed across the spectrum. Experiments on the WildSVDD dataset demonstrate that high-frequency subbands provide essential complementary cues. Our framework significantly outperforms 16 kHz-sampled models, proving that high-resolution audio and strategic subband integration are critical for robust in-the-wild detection.

歌声伪造高分辨音频子频段建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。