用人类可察觉的语言特征提升伪造音频检测能力
Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
- 结合人类可听的语言特征构建因果模型
- 发现语音特征与伪造标签间存在显著因果关联
- 适合研究语音安全与人机协同的AI工程师
多种伪造音频(如模仿、重放攻击、深度伪造)对信息真实性构成社会挑战。近期,研究人员联合语言学专家,为伪造音频样本标注了人类可听的语言特征(EDLFs):音高、停顿、辅音爆破起始/结束冲激、呼吸声及整体音质。实验表明,将这些特征融入传统音频特征后,多个深度伪造检测算法性能得到提升。本文基于包含多类伪造音频并附有人类语言学标注的混合数据集,研究了可察觉语言特征与音频标签之间的因果关系,并与专家标注结果对比验证。结果表明,因果模型能有效识别语言特征在辨别伪造音频中的作用,凸显将人类知识融入AI模型的必要性与潜力。该方法可作为训练人类辨别伪造音频的基础,也可自动化标注EDLF以提升现有AI检测器性能。
原文摘要 · Abstract (English)
Several types of spoofed audio, such as mimicry, replay attacks, and deepfakes, have created societal challenges to information integrity. Recently, researchers have worked with sociolinguistics experts to label spoofed audio samples with Expert Defined Linguistic Features (EDLFs) that can be discerned by the human ear: pitch, pause, word-initial and word-final release bursts of consonant stops, audible intake or outtake of breath, and overall audio quality. It is established that there is an improvement in several deepfake detection algorithms when they augmented the traditional and common features of audio data with these EDLFs. In this paper, using a hybrid dataset comprised of multiple types of spoofed audio augmented with sociolinguistic annotations, we investigate causal discovery and inferences between the discernible linguistic features and the label in the audio clips, comparing the findings of the causal models with the expert ground truth validation labeling process. Our findings suggest that the causal models indicate the utility of incorporating linguistic features to help discern spoofed audio, as well as the overall need and opportunity to incorporate human knowledge into models and techniques for strengthening AI models. The causal discovery and inference can be used as a foundation of training humans to discern spoofed audio as well as automating EDLFs labeling for the purpose of performance improvement of the common AI-based spoofed audio detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。