arXiv:2410.15577cs.SDeess.AS2024-10被引 2

用自动标注语言特征提升伪造语音检测效果。

ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection

  • 基于社会语言学专家提取的语言特征,训练自动标注框架ALDAS。
  • 使用自动生成的语言特征,检测性能优于仅用声学特征的模型。
  • 专家验证了自动生成标签的有效性,适合语音安全与深度伪造研究者。

伪造语音(即被篡改或由AI生成的深度伪造音频)仅依靠声学特征难以检测。近期一些创新工作通过人工标注英语语音的音位和音系特征,显著提升了检测模型性能。然而,人工标注存在可扩展性问题,促使研究者探索自动标注方案。本文提出一种名为ALDAS(Audio-Linguistic Data Augmentation for Spoofed audio detection)的AI框架,实现语言特征的自动标注。该框架在社会语言学专家选定并提取的语言特征上进行训练,生成的自标注特征用于评估预测质量。结果显示,尽管性能提升不及完全依赖真实标注特征的情况,但相比仅使用声学特征的模型仍有明显改善;同时,专家对ALDAS生成的标签进行了验证,确认其有效性。

原文摘要 · Abstract (English)

Spoofed audio, i.e. audio that is manipulated or AI-generated deepfake audio, is difficult to detect when only using acoustic features. Some recent innovative work involving AI-spoofed audio detection models augmented with phonetic and phonological features of spoken English, manually annotated by experts, led to improved model performance. While this augmented model produced substantial improvements over traditional acoustic features based models, a scalability challenge motivates inquiry into auto labeling of features. In this paper we propose an AI framework, Audio-Linguistic Data Augmentation for Spoofed audio detection (ALDAS), for auto labeling linguistic features. ALDAS is trained on linguistic features selected and extracted by sociolinguistics experts; these auto labeled features are used to evaluate the quality of ALDAS predictions. Findings indicate that while the detection enhancement is not as substantial as when involving the pure ground truth linguistic features, there is improvement in performance while achieving auto labeling. Labels generated by ALDAS are also validated by the sociolinguistics experts.

语音检测深度伪造自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。