arXiv:2512.04175cs.CV2025-12被引 1

通过模拟面部运动不一致,提升深伪视频检测的泛化能力。

Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection

  • 用运动基分解人脸关键点,人为制造自然动作违例。
  • 在多个基准上实现当前最优的跨域检测效果。
  • 适合关注视频伪造检测泛化性的研究者和安全团队。

将深伪检测泛化至未见篡改方法仍是关键挑战。现有方法通过手工生成伪影训练网络以提取更通用线索,虽在静态图像中有效,但扩展到视频领域仍存难题。现有方法将时间伪影建模为帧间不稳定性,忽略了关键漏洞:不同面部区域间自然运动依赖关系的破坏。本文提出一种合成视频生成方法,创建带有细微运动不一致的训练数据。训练自编码器将人脸关键点配置分解为运动基,通过操控这些基,有选择性地破坏面部运动的自然相关性,并通过人脸形变将伪影引入原始视频。在该数据上训练的网络能识别此类复杂生物力学缺陷,在多个主流基准上取得最先进的泛化性能。

原文摘要 · Abstract (English)

Generalizing deepfake detection to unseen manipulations remains a key challenge. A recent approach to tackle this issue is to train a network with pristine face images that have been manipulated with hand-crafted artifacts to extract more generalizable clues. While effective for static images, extending this to the video domain is an open issue. Existing methods model temporal artifacts as frame-to-frame instabilities, overlooking a key vulnerability: the violation of natural motion dependencies between different facial regions. In this paper, we propose a synthetic video generation method that creates training data with subtle kinematic inconsistencies. We train an autoencoder to decompose facial landmark configurations into motion bases. By manipulating these bases, we selectively break the natural correlations in facial movements and introduce these artifacts into pristine videos via face morphing. A network trained on our data learns to spot these sophisticated biomechanical flaws, achieving state-of-the-art generalization results on several popular benchmarks.

视频生成深度伪造检测泛化运动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。