构建含语言特征的语音数据集,提升伪造语音检测能力。
Linguistically Augmented Audio Speech Data (LinguAS)

- 引入五类专家定义的语言特征,增强语音检测模型的上下文理解。
- 使用该数据集训练的模型性能显著优于ASVspoof 2021基线和HuBert等SSL模型。
- 适合语音安全、深度伪造检测方向的研究者使用。
恶意生成的虚假语音(如深度伪造和语音欺骗)正以惊人速度扩散,而现有检测模型大多仅依赖帧级音频特征,忽视了长时间尺度下的语言线索。为弥补这一差距,我们提出Linguistically Augmented Audio Speech Data (LinguAS),一个包含超过800个真实与深度伪造语音样本的数据集,每条样本均标注了五类专家定义的语言特征(EDLFs),这些特征在英语口语中常见且具有自然人类语音的典型性。数据集涵盖四类均衡的语音欺骗攻击类型,并按比例分配真实语音样本;同时提供说话人性别及每段伪造语音的生成器/来源元数据,便于更精细的模型训练。实验表明,基于包含EDLFs增强数据训练的模型,在性能上显著超越ASVspoof 2021深度学习基线以及HuBert和XLSR等自监督模型。LinguAS通过融合语言、性别和生成器元数据,为语音深度伪造研究提供了强调真实语言特性的高质量数据资源。数据与代码已公开。
原文摘要 · Abstract (English)
Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve. Yet, most detection models are trained to make inference on frame-level audio features alone without leveraging valuable linguistic cues at larger timescales. To address this gap, we present Linguistically Augmented Audio Speech Data (LinguAS), a dataset of genuine and deepfaked audio samples annotated with five strategically-chosen, Expert-Defined Linguistic Features (EDLFs) that occur frequently in spoken English and are characteristic of natural human speech. LinguAS contains over 800 audio samples, each of which are annotated with EDLFs. The dataset has a balanced number of four spoofed audio attack types and a proportionate number of genuine speech samples. We also include metadata on speaker gender and the generator/source for each spoofed audio sample, offering more granularity for model training. We found that models trained on data augmented with EDLFs had improved model performance significantly beyond the ASVspoof 2021 deep learning baselines and SSL models like HuBert and XLSR. LinguAS's augmented linguistic, gender, and generator metadata provide audio deepfake researchers with a dataset that emphasizes real human language traits to improve model inference of faked speech. Data and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。