提出多说话人对话音频伪造分类体系并发布首个相关数据集
Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
- 按部分或全部替换说话人,构建多说话人音频伪造分类框架
- 发布2830段真实与合成对话数据,涵盖不同性别和自然度
- 为对话场景伪造检测提供基准,适合安全与可信音频研究者
文本转语音(TTS)技术的快速发展使音频深度伪造日益逼真且易获取,引发严重的安全与信任问题。现有研究多聚焦单说话人音频伪造检测,但现实中多说话人对话场景的恶意应用正成为亟待关注却未被充分探索的威胁。为此,本文提出多说话人对话音频伪造的概念性分类体系,区分部分篡改(一个或多个说话人被修改)与完全合成(整段对话生成)。作为初步工作,我们构建了多说话人对话音频深度伪造数据集(MsCADD),包含2830个音频片段,涵盖真实与完全合成的双说话人对话,使用VITS和SoundStorm-based NotebookLM模型生成,模拟不同说话人性别及对话自然度变化。该数据集仅包含TTS类伪造。我们在该数据集上对三种神经基线模型(LFCC-LCNN、RawNet2、Wav2Vec 2.0)进行基准测试,报告了F1分数、准确率、真正例率(TPR)和真负例率(TNR)。结果表明,这些模型提供了有效基准,但同时也凸显出在复杂对话动态下可靠检测合成语音仍存在显著差距。本数据集与基准为未来对话场景深度伪造检测研究奠定基础,该领域虽严重不足,却是保障音频信息可信性的关键挑战。MsCADD数据集已公开,以支持可复现性与社区基准测试。
原文摘要 · Abstract (English)
The rapid advances in text-to-speech (TTS) technologies have made audio deepfakes increasingly realistic and accessible, raising significant security and trust concerns. While existing research has largely focused on detecting single-speaker audio deepfakes, real-world malicious applications with multi-speaker conversational settings is also emerging as a major underexplored threat. To address this gap, we propose a conceptual taxonomy of multi-speaker conversational audio deepfakes, distinguishing between partial manipulations (one or multiple speakers altered) and full manipulations (entire conversations synthesized). As a first step, we introduce a new Multi-speaker Conversational Audio Deepfakes Dataset (MsCADD) of 2,830 audio clips containing real and fully synthetic two-speaker conversations, generated using VITS and SoundStorm-based NotebookLM models to simulate natural dialogue with variations in speaker gender, and conversational spontaneity. MsCADD is limited to text-to-speech (TTS) types of deepfake. We benchmark three neural baseline models; LFCC-LCNN, RawNet2, and Wav2Vec 2.0 on this dataset and report performance in terms of F1 score, accuracy, true positive rate (TPR), and true negative rate (TNR). Results show that these baseline models provided a useful benchmark, however, the results also highlight that there is a significant gap in multi-speaker deepfake research in reliably detecting synthetic voices under varied conversational dynamics. Our dataset and benchmarks provide a foundation for future research on deepfake detection in conversational scenarios, which is a highly underexplored area of research but also a major area of threat to trustworthy information in audio settings. The MsCADD dataset is publicly available to support reproducibility and benchmarking by the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。