一张网检测多种视频伪造,能同时识别深伪、修补、拼接等造假手法。
MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence
- 融合多维度取证特征,跨空间与时间尺度分析视频异常。
- 在混合篡改场景下性能领先,单类检测也接近专用模型水平。
- 适合需要通用视频真伪检测的场景,如内容审核与司法取证。
尽管视频可被多种方式伪造,但现有取证网络通常仅针对单一篡改类型(如深伪、修补)。这带来实际问题:伪造手法事先未知。为此,我们提出MVFNet——一种可检测多种篡改类型的通用视频取证网络,涵盖修补、深伪、拼接和编辑。该网络通过提取并联合分析多种取证特征模态,捕捉视频中的空间与时间异常。为可靠检测任意形状和大小的虚假内容,引入新型多尺度分层变压器模块,实现跨尺度的取证不一致识别。实验表明,该网络在多种篡改共存的一般场景中达到当前最优表现,在特定场景下亦可媲美专用检测器。
原文摘要 · Abstract (English)
While videos can be falsified in many different ways, most existing forensic networks are specialized to detect only a single manipulation type (e.g. deepfake, inpainting). This poses a significant issue as the manipulation used to falsify a video is not known a priori. To address this problem, we propose MVFNet - a multipurpose video forensics network capable of detecting multiple types of manipulations including inpainting, deepfakes, splicing, and editing. Our network does this by extracting and jointly analyzing a broad set of forensic feature modalities that capture both spatial and temporal anomalies in falsified videos. To reliably detect and localize fake content of all shapes and sizes, our network employs a novel Multi-Scale Hierarchical Transformer module to identify forensic inconsistencies across multiple spatial scales. Experimental results show that our network obtains state-of-the-art performance in general scenarios where multiple different manipulations are possible, and rivals specialized detectors in targeted scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。