arXiv:2511.18993cs.CV2025-11被引 5

通过跨模态语音重建定位深度伪造视频的时间片段。

AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization

  • 利用音视频互换重建语音表征,检测伪造区域。
  • 在多个数据集上准确率提升超8.9个百分点。
  • 适合需要精准识别伪造片段的媒体安全研究者。

随着高精度音视频合成技术的快速发展,如用于隐蔽恶意篡改的内容,保障数字媒体完整性变得至关重要。本文提出一种基于音频-视觉语音表征重建(AuViRe)的深度伪造时间定位新方法。具体而言,该方法基于一模态(如唇动)重建另一模态(如音频波形)的语音表征。在伪造视频片段中,跨模态重建难度显著增加,导致差异放大,从而提供鲁棒的判别性线索,实现精确的时间伪造定位。AuViRe 在 LAV-DF 上比现有最优方法提升 +8.9 [email protected],AV-Deepfake1M 上提升 +9.6 [email protected],且在真实场景实验中提升 +5.1 AUC。代码已开源:https://github.com/mever-team/auvire。

原文摘要 · Abstract (English)

With the rapid advancement of sophisticated synthetic audio-visual content, e.g., for subtle malicious manipulations, ensuring the integrity of digital media has become paramount. This work presents a novel approach to temporal localization of deepfakes by leveraging Audio-Visual Speech Representation Reconstruction (AuViRe). Specifically, our approach reconstructs speech representations from one modality (e.g., lip movements) based on the other (e.g., audio waveform). Cross-modal reconstruction is significantly more challenging in manipulated video segments, leading to amplified discrepancies, thereby providing robust discriminative cues for precise temporal forgery localization. AuViRe outperforms the state of the art by +8.9 [email protected] on LAV-DF, +9.6 [email protected] on AV-Deepfake1M, and +5.1 AUC on an in-the-wild experiment. Code available at https://github.com/mever-team/auvire.

深度伪造音视频分析时间定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。