通过双流图学习与解耦,精准定位视频篡改片段。
DDNet: A Dual-Stream Graph Learning and Disentanglement Framework for Temporal Forgery Localization

- 双流设计:分别捕捉局部异常与长程语义关联
- 在ForgeryNet和TVIL上[email protected]提升9%,跨域鲁棒性更强
- 适合视频安全检测、深度伪造分析的研究者
AIGC技术快速发展,使得仅篡改视频中少量片段即可误导观众,导致视频级检测失效。因此,时间篡改定位(TFL)——精确定位被篡改区域——变得至关重要。然而,现有方法多受限于局部视角,难以捕捉全局异常。为此,本文提出双流图学习与解耦框架DDNet。通过时序距离流捕获局部伪影,语义内容流建模长程依赖,避免全局线索被局部平滑性淹没。进一步引入痕迹解耦与自适应(TDA)以分离通用篡改指纹,并采用跨层级特征嵌入(CLFE)融合多层特征构建鲁棒表征。在ForgeryNet与TVIL基准上的实验表明,本方法在[email protected]上较当前最优方法提升约9%,且跨域泛化能力显著增强。
原文摘要 · Abstract (English)
The rapid evolution of AIGC technology enables misleading viewers by tampering mere small segments within a video, rendering video-level detection inaccurate and unpersuasive. Consequently, temporal forgery localization (TFL), which aims to precisely pinpoint tampered segments, becomes critical. However, existing methods are often constrained by \emph{local view}, failing to capture global anomalies. To address this, we propose a \underline{d}ual-stream graph learning and \underline{d}isentanglement framework for temporal forgery localization (DDNet). By coordinating a \emph{Temporal Distance Stream} for local artifacts and a \emph{Semantic Content Stream} for long-range connections, DDNet prevents global cues from being drowned out by local smoothness. Furthermore, we introduce Trace Disentanglement and Adaptation (TDA) to isolate generic forgery fingerprints, alongside Cross-Level Feature Embedding (CLFE) to construct a robust feature foundation via deep fusion of hierarchical features. Experiments on ForgeryNet and TVIL benchmarks demonstrate that our method outperforms state-of-the-art approaches by approximately 9\% in [email protected], with significant improvements in cross-domain robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。