arXiv:2604.12650cs.CVcs.MM2026-04

提出监听态伪造检测新任务,破解交互场景中欺骗性深伪漏洞。

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

  • 构建首个监听态伪造数据集ListenForge,涵盖五种生成方法。
  • 提出MANet模型,融合音频语义与微动一致性分析,准确率显著提升。
  • 揭示现有说话态检测模型在监听场景失效,适合多模态安全研究者。

现有深度伪造检测研究主要聚焦于说话状态下的伪造,即通过改变说话人外观或声音生成虚假内容。然而,在真实交互场景中,攻击者常交替伪造说话与倾听状态,以增强情境的逼真度和说服力。尽管监听态伪造检测尚未被充分探索,且受限于数据集和方法的缺乏,但合成倾听反应的相对低质量为当前检测技术提供了突破口。本文首次提出监听态深伪检测(LDD)任务,构建了首个专为此任务设计的数据集ListenForge,采用五种倾听头生成(LHG)方法创建。针对倾听伪造的独特特征,提出MANet模型,该模型通过捕捉观众视频中的细微动作不一致,并利用说话人音频语义引导跨模态融合。大量实验表明,现有说话态检测模型在监听场景表现不佳,而MANet在ListenForge上取得显著更优性能。本工作强调需突破传统以说话为中心的检测范式,为交互通信中的多模态伪造分析开辟新方向。数据集与代码已公开于https://anonymous.4open.science/r/LDD-B4CB。

原文摘要 · Abstract (English)

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in realistic interaction settings, attackers often alternate between falsifying speaking and listening states to mislead their targets, thereby enhancing the realism and persuasiveness of the scenario. Although the detection of 'listening deepfakes' remains largely unexplored and is hindered by a scarcity of both datasets and methodologies, the relatively limited quality of synthesized listening reactions presents an excellent breakthrough opportunity for current deepfake detection efforts. In this paper, we present the task of Listening Deepfake Detection (LDD). We introduce ListenForge, the first dataset specifically designed for this task, constructed using five Listening Head Generation (LHG) methods. To address the distinctive characteristics of listening forgeries, we propose MANet, a Motion-aware and Audio-guided Network that captures subtle motion inconsistencies in listener videos while leveraging speaker's audio semantics to guide cross-modal fusion. Extensive experiments demonstrate that existing Speaking Deepfake Detection (SDD) models perform poorly in listening scenarios. In contrast, MANet achieves significantly superior performance on ListenForge. Our work highlights the necessity of rethinking deepfake detection beyond the traditional speaking-centric paradigm and opens new directions for multimodal forgery analysis in interactive communication settings. The dataset and code are available at https://anonymous.4open.science/r/LDD-B4CB.

深伪检测多模态监听伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。