用图像监督提升视频部分伪造检测能力,无需额外计算开销
Learning Unified Video and Image Representation for Video Face Forgery Detection

- 统一编码器+多任务学习,联合建模视频与图像特征
- 伪标签机制对齐视频帧与静态图像表示,减少分布差异
- 在基准数据集上优于现有方法,特别擅长检测部分伪造视频
人脸伪造检测对保障面部数据安全至关重要,因人脸操纵技术和生成模型快速发展。现有视频伪造检测方法通常假设所有帧均被篡改,而对仅部分帧被修改的视频检测仍具挑战。为此,我们提出新框架UVIF,利用额外标注图像提供细粒度监督,以检测视频中的局部伪造。UVIF采用统一编码器和多任务学习范式,联合建模视频与图像特征。使用带时间融合模块的2D主干网络作为统一编码器,设计伪标签过程对齐视频帧与静态图像表征,并引入面向视频的特征对齐策略,缩小视频与图像间的分布差距。在多个基准数据集上的大量实验表明,该框架在检测部分伪造视频方面显著优于现有方法,且无额外计算开销。
原文摘要 · Abstract (English)
Face forgery detection is crucial for preserving the security and integrity of facial data given the rapid developments in face manipulation techniques and deep generative models. Existing methods for video face forgery detection typically assume that all frames in a forged video are manipulated, while detecting partially forged videos that contain only a subset of altered frames remains challenging. To address this issue, we propose a novel framework, UVIF, that utilizes additional annotated images to provide fine-grained supervision for detecting partial forgeries in videos. UVIF employs a unified encoder and a multi-task learning paradigm to jointly model facial videos and images for boosted video face forgery detection. A 2D backbone with temporal fusion modules is employed as the unified encoder. A pseudo labeling process is designed for video frames to bridge their representations with those of static images. A video-oriented feature alignment strategy is further introduced to reduce the distribution gap between videos and images. Extensive experiments on benchmark datasets demonstrate the effectiveness of our framework, which outperforms state-of-theart methods in detecting partially forged videos while introducing no additional computational overhead. Our code is available at https://github.com/haotianll/UVIF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。