arXiv:2507.16220cs.SDcs.CR2025-07中稿 · IEEE International…被引 2

提出真实场景下长时语音伪造检测与定位新方案

LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech

  • 构建可控的长时、多说话人、带噪语音生成流程
  • 基于自监督前端+简单后端,检测性能显著提升
  • 适用于真实复杂环境下的伪造语音分析与定位

本文提出LENS-DF,一种面向复杂现实音频条件下的语音深度伪造检测与时间定位综合训练与评估方法。其生成部分可可控地输出具有较长时长、噪声环境及多说话人特征的音频。对应的检测与定位协议采用模型进行评估。实验基于自监督学习前端与简单后端开展,结果表明,使用LENS-DF生成数据训练的模型在各项指标上均优于传统方法,验证了该方案在鲁棒性语音伪造检测与定位中的有效性。此外,通过消融实验考察了各变量的影响,分析其对实际挑战的适配性。

原文摘要 · Abstract (English)

This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs audios from the input dataset with several critical characteristics, such as longer duration, noisy conditions, and containing multiple speakers, in a controllable fashion. The corresponding detection and localization protocol uses models. We conduct experiments based on self-supervised learning front-end and simple back-end. The results indicate that models trained using data generated with LENS-DF consistently outperform those trained via conventional recipes, demonstrating the effectiveness and usefulness of LENS-DF for robust audio deepfake detection and localization. We also conduct ablation studies on the variations introduced, investigating their impact on and relevance to realistic challenges in the field.

语音伪造检测时间定位自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。