arXiv:2411.05969cs.SDcs.CL2024-11

跨学科融合语言学与AI,提升语音深度伪造检测能力

Toward Transdisciplinary Approaches to Audio Deepfake Discernment

  • 结合语言学知识改进AI模型对语音变异的理解
  • 推动专家参与的智能检测系统构建
  • 适合语音安全、AI伦理领域研究者参考

本文呼吁跨学科合作,以应对语音深度伪造检测的挑战。随着生成逼真假语音的工具激增,检测技术却相对滞后。当前AI模型因缺乏对语言内在变异性及人类语音复杂性的全面理解,难以有效识别深度伪造。近期跨学科研究将语言学知识融入AI方法,为构建‘专家在环’系统提供路径,推动从‘专家无关’的通用模型向更鲁棒、更全面的检测方法演进。

原文摘要 · Abstract (English)

This perspective calls for scholars across disciplines to address the challenge of audio deepfake detection and discernment through an interdisciplinary lens across Artificial Intelligence methods and linguistics. With an avalanche of tools for the generation of realistic-sounding fake speech on one side, the detection of deepfakes is lagging on the other. Particularly hindering audio deepfake detection is the fact that current AI models lack a full understanding of the inherent variability of language and the complexities and uniqueness of human speech. We see the promising potential in recent transdisciplinary work that incorporates linguistic knowledge into AI approaches to provide pathways for expert-in-the-loop and to move beyond expert agnostic AI-based methods for more robust and comprehensive deepfake detection.

语音伪造跨学科AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。