arXiv:2609.03446cs.CV2026-09

针对视频伪造检测的持续学习,分离保留时空特征以提升适应能力。

Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

论文配图:Preserving Knowledge across Space and Time for Continual Video Deepfake Detection
图 1 · 摘自论文原文
  • 在频域分解视频特征为时空、时序、空域三模态,分别独立保存。
  • 跨模态正交损失使多模态表示不冗余,提升检测鲁棒性。
  • 适用于需要长期更新的视频伪造检测场景,尤其适合新伪造类型泛化。

持续出现的高质量视频深度伪造要求检测器能不断适应新伪造模式,但现有方法多针对图像设计,难以捕捉视频特有的时空线索。与仅含空间伪影的图像不同,视频伪造在空间和时间轴上均留下显著证据,因此在连续模型更新中需分别保留各模态信息。为此,我们提出一种持续视频伪造检测框架——模态特定频域蒸馏(MSFD),在频域将视频特征显式分解为空间、时间及时空模态,实现各模态独立保存。由于不同伪造类型对空间与时间线索的依赖程度不同,该机制可有效适应多变任务。此外,MSFD引入跨模态去相关损失,促使时空表示与单模态线索保持正交。大量实验表明,该框架在多种持续视频伪造检测场景下,相比当前最优方法具有更强的适应能力与更优的性能保持效果。

原文摘要 · Abstract (English)

The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.

视频伪造持续学习频域蒸馏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。