arXiv:2511.10212cs.CV2025-11被引 1

通过预测下一帧特征提升伪造视频检测与定位精度

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

  • 单阶段训练中引入跨模态与单模态的帧预测机制
  • 在多个数据集上实现强泛化能力与精准时间定位
  • 适合需要高精度局部伪造检测的应用场景

近期多模态深度伪造检测方法认为,单阶段监督训练难以在未见篡改手法和数据集上泛化。然而,这些追求泛化的方案通常需在真实样本上预训练。此外,现有方法主要关注音视频不一致,可能忽略导致该现象的模内伪影,因而对保持音视频对齐的篡改手段失效。为此,我们提出一种单阶段训练框架,通过引入单模态与跨模态特征的下一帧预测,增强模型泛化能力。同时,我们设计窗口级注意力机制,捕捉预测帧与实际帧间的差异,使模型能检测每帧附近的局部伪影,这对完整伪造视频的分类及部分伪造样本的精确时空定位至关重要。模型在多个基准数据集上评估,展现出优异的泛化性能与精准的时间定位能力。

原文摘要 · Abstract (English)

Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and datasets. However, such approaches that target generalization require pretraining over real samples. Additionally, these methods primarily focus on detecting audio-visual inconsistencies and may overlook intra-modal artifacts causing them to fail against manipulations that preserve audio-visual alignment. To address these limitations, we propose a single-stage training framework that enhances generalization by incorporating next-frame prediction for both uni-modal and cross-modal features. Additionally, we introduce a window-level attention mechanism to capture discrepancies between predicted and actual frames, enabling the model to detect local artifacts around every frame, which is crucial for accurately classifying fully manipulated videos and effectively localizing deepfake segments in partially spoofed samples. Our model, evaluated on multiple benchmark datasets, demonstrates strong generalization and precise temporal localization.

深度伪造检测时序定位多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。