研究发现伴奏主要起数据增强作用,而非提供节奏等内在线索。
How Does Instrumental Music Help SingFake Detection?
- 通过测试不同模型和频段,分析伴奏影响机制
- 微调后模型更依赖说话人浅层特征,忽略语义信息
- 适合关注语音伪造检测可解释性的研究人员
尽管已有多种模型用于检测歌唱语音深度伪造(SingFake),但其在伴奏影响下的工作原理仍不明确。本文从行为效应和表征效应两方面展开研究:测试不同主干网络、无配对伴奏音轨及频率子带;探测微调如何改变编码器的语音与音乐表征能力。结果表明,伴奏主要起数据增强作用,而非提供节奏或和声等内在线索;微调使模型更依赖浅层说话人特征,降低对内容、副语言和语义信息的敏感性。这些发现揭示了模型如何利用人声与伴奏线索,有助于设计更具可解释性和鲁棒性的SingFake检测系统。
原文摘要 · Abstract (English)
Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different backbones, unpaired instrumental tracks, and frequency subbands. To analyze the representational effect, we probe how fine-tuning alters encoders' speech and music capabilities. Our results show that instrumental accompaniment acts mainly as data augmentation rather than providing intrinsic cues (e.g., rhythm or harmony). Furthermore, fine-tuning increases reliance on shallow speaker features while reducing sensitivity to content, paralinguistic, and semantic information. These insights clarify how models exploit vocal versus instrumental cues and can inform the design of more interpretable and robust SingFake detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。