针对印尼语和泰语非母语语音伪造,提出更有效的检测方法。
Detecting Spoof Voices in Asian Non-Native Speech: An Indonesian and Thai Case Study
- 用梅尔倒谱、线性谱倒谱等特征训练分类器。
- 结合非母语数据训练的模型检测效果提升明显。
- 适合多语言语音安全研究者参考。
本研究聚焦于非母语语音的欺骗检测对策(CMs),特别针对印尼语和泰语使用者。我们构建了一个包含母语与非母语语音的数据集以支持研究。从语音数据中提取了三种关键特征(MFCC、LFCC 和 CQCC),并采用三种经典机器学习分类器(CatBoost、XGBoost、GMM)基于母语及混合(母语与非母语)语音数据开发鲁棒的欺骗检测系统,形成两类对策:母语型与混合型。对这两类对策在母语与非母语语音数据集上的表现进行了评估。研究发现,仅使用母语数据训练的检测系统在处理非母语语音时存在显著挑战,凸显了领域定制解决方案的必要性。所提方法展现出更强的检测能力,证明将非母语语音数据纳入训练过程的重要性。该工作为多样化语言场景下的有效欺骗检测系统奠定了基础。
原文摘要 · Abstract (English)
This study focuses on building effective spoofing countermeasures (CMs) for non-native speech, specifically targeting Indonesian and Thai speakers. We constructed a dataset comprising both native and non-native speech to facilitate our research. Three key features (MFCC, LFCC, and CQCC) were extracted from the speech data, and three classic machine learning-based classifiers (CatBoost, XGBoost, and GMM) were employed to develop robust spoofing detection systems using the native and combined (native and non-native) speech data. This resulted in two types of CMs: Native and Combined. The performance of these CMs was evaluated on both native and non-native speech datasets. Our findings reveal significant challenges faced by Native CM in handling non-native speech, highlighting the necessity for domain-specific solutions. The proposed method shows improved detection capabilities, demonstrating the importance of incorporating non-native speech data into the training process. This work lays the foundation for more effective spoofing detection systems in diverse linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。