arXiv:2606.22772cs.CV2026-06中稿 · the IEEE Internati…

通过反事实帧一致性检测唇同步深度伪造,定位精准且泛化强。

LoCC: Detection and Localization of Lip-Syncing Deepfakes via Counterfactual Frame Consistency

论文配图:LoCC: Detection and Localization of Lip-Syncing Deepfakes via Counterfactual Frame Consistency
图 1 · 摘自论文原文
  • 基于时间邻近帧生成反事实估计,检验每帧一致性
  • 在多个数据集上准确率超现有方法,跨压缩等级表现稳定
  • 适合需要高精度伪造检测的视频安全与内容审核场景

唇同步深度伪造是极具挑战性的篡改媒体形式,其伪影主要集中于嘴部区域且随时间动态变化。检测此类伪造需精确建模唇部运动的时间与空间特性。本文提出LoCC框架,实现细粒度的段级与帧级检测与定位。不同于以往整体分析视频的方法,本方法通过评估每帧是否与来自其时间邻近帧的反事实估计一致来判断真伪。真实视频表现出强而稳定的连续性,而唇同步伪造则引入局部不一致性。采用教师-学生学习范式,模型有效捕捉帧级差异,在多个基准数据集(包括LAV-DF、AVDF1M、FakeAVCeleb和KODF)上优于当前最优方法,并在不同压缩级别和数据集间具有良好泛化能力。

原文摘要 · Abstract (English)

Lip-syncing deepfakes are among the most challenging forms of manipulated media because their artifacts are localized almost exclusively to the mouth region and evolve dynamically over time. Detecting such deepfakes requires precise temporal and spatial modeling of lip motion. In this paper, we propose LoCC, a novel detection framework that performs fine-grained detection and localization of lip-syncing deepfakes at both segment and frame levels. Unlike prior approaches that analyze videos holistically, our method evaluates whether each frame aligns with a counterfactual estimate generated from its temporal neighbors. Real videos exhibit strong and stable consistency, whereas lip-sync deepfakes introduce localized inconsistencies. Following a teacher-student learning paradigm, our model effectively captures these frame-level discrepancies and achieves superior performance over state-of-the-art methods on multiple benchmark lip-syncing deepfake datasets, including LAV-DF, AVDF1M, FakeAVCeleb, and KODF, and generalizes well across compression levels and datasets.

深度伪造检测唇同步伪造反事实分析视频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。