通过重建误差定位深层伪造片段,无需逐帧标注。
Mining Forgery Traces from Reconstruction Error: A Weakly Supervised Framework for Multimodal Deepfake Temporal Localization
- 用真实视频训练掩码自编码器,利用重建误差识别伪造段。
- 在LAV-DF数据集上达到当前最优弱监督定位性能。
- 适合关注伪造检测与视频安全的开发者和研究人员。
现代深层伪造已演变为局部且间歇性篡改,需细粒度时间定位以应对严重数字安全风险。帧级标注成本高昂,弱监督方法(仅依赖视频级标签)成为实际必需。为此,我们提出基于重建误差的弱监督时间伪造定位框架RT-DeepLoc,通过重建误差识别伪造内容。该框架使用仅在真实数据上训练的掩码自编码器(MAE)学习内在时空模式,使伪造片段产生显著重建偏差,从而提供缺失的细粒度线索,实现无需密集人工标注的精准定位。为稳健利用这些指标,引入新型非对称视频内对比损失(AICL),通过重建提示引导真实特征紧凑性,建立稳定决策边界,增强局部判别力,同时保持对先进生成模型新伪造的泛化能力。在大规模数据集(包括LAV-DF)上的实验表明,RT-DeepLoc在弱监督时间伪造定位任务中达到领先性能。
原文摘要 · Abstract (English)
Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digital security risks. The prohibitive cost of frame-level annotation makes weakly supervised methods a practical necessity, which rely only on video-level labels. To this end, we propose Reconstruction-based Temporal Deepfake Localization (RT-DeepLoc), a weakly supervised temporal forgery localization framework that identifies forgeries via reconstruction errors. Our framework uses a Masked Autoencoder (MAE) trained exclusively on authentic data to learn its intrinsic spatiotemporal patterns; this allows the model to produce significant reconstruction discrepancies for forged segments, effectively providing the missing fine-grained cues for accurate localization without demanding dense human annotations. To robustly leverage these indicators, we introduce a novel Asymmetric Intra-video Contrastive Loss (AICL). By focusing on the compactness of authentic features guided by these reconstruction cues, AICL establishes a stable decision boundary that enhances local discrimination while preserving generalization to unseen forgeries by advanced generative models. Extensive experiments on large-scale datasets, including LAV-DF, demonstrate that RT-DeepLoc achieves state-of-the-art performance in weakly-supervised temporal forgery localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。