arXiv:2604.19949eess.AS2026-04ACL被引 1

首个印地语语音深度伪造检测数据集,提出新模型有效识别合成语音。

Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages

论文配图:Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
图 1 · 摘自论文原文
  • 构建首个印地语语音深度伪造数据集ICF,涵盖多语言、多声线和多种编码器。
  • 现有主流检测模型在印地语上表现差,因语音音韵与语调差异大。
  • 提出超球面模型SATYAM,融合语义与韵律特征,提升跨语言检测能力。

音频大模型(ALMs)的发展推动了神经音频编码器(NACs)的应用,催生出高度逼真的语音深度伪造(称作CodecFakes, CFs)。然而,现有研究主要集中在英语或中文,印地语等印欧语言的脆弱性仍被忽视。为此,我们构建了首个大规模基准数据集Indic-CodecFake(ICF),包含多印地语、多说话人及多种NAC生成的真实与合成语音。采用IndicSUPERB作为真实语音来源生成ICF。实验表明,基于英语数据训练的先进检测模型在ICF上泛化能力差,凸显印地语语音在音韵与语调上的多样性挑战。进一步评估了主流ALMs在零样本设置下的表现,发现其性能普遍不佳。为此,我们提出SATYAM——一种专为印地语语音深度伪造检测设计的超球面音频大模型。SATYAM通过双阶段融合:利用Whisper提取语义表示,结合TRILLsson提取韵律特征,并在超球面空间中使用巴塔查里亚距离进行对齐,最终实现语音与文本提示的联合建模。该架构有效捕捉语音内部(语义-韵律)及跨模态(语音-文本)的层级关系。大量实验证明,SATYAM在ICF基准上持续优于现有端到端与基于ALM的基线方法。

原文摘要 · Abstract (English)

The rapid advancement of Audio Large Language Models (ALMs), driven by Neural Audio Codecs (NACs), has led to the emergence of highly realistic speech deepfakes, commonly referred to as CodecFakes (CFs). Consequently, CF detection has attracted increasing attention from the research community. However, existing studies predominantly focus on English or Chinese, leaving the vulnerability of Indic languages largely unexplored. To bridge this gap, we introduce Indic-CodecFake (ICF) dataset, the first large-scale benchmark comprising real and NAC-synthesized speech across multiple Indic languages, diverse speaker profiles, and multiple NAC types. We use IndicSUPERB as the real speech corpus for generation of ICF dataset. Our experiments demonstrate that state-of-the-art (SOTA) CF detectors trained on English-centric datasets fail to generalize to ICF, underscoring the challenges posed by phonetic diversity and prosodic variability in Indic speech. Further, we present systematic evaluation of SOTA ALMs in a zero-shot setting on ICF dataset. We evaluate these ALMs as they have shown effectiveness for different speech tasks. However, our findings reveal that current ALMs exhibit consistently poor performance. To address this, we propose SATYAM, a novel hyperbolic ALM tailored for CF detection in Indic languages. SATYAM integrates semantic representations from Whisper and prosodic representations from TRILLsson using through Bhattacharya distance in hyperbolic space and subsequently performs the same alignment procedure between the fused speech representation and an input conditioning prompt. This dual-stage fusion framework enables SATYAM to effectively model hierarchical relationships both within speech (semantic-prosodic) and across modalities (speech-text). Extensive evaluations show that SATYAM consistently outperforms competitive end-to-end and ALM-based baselines on the ICF benchmark.

语音伪造深度伪造检测印地语超球面模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。