提出统一上下文感知对比学习框架,精准定位视频中局部篡改片段。
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
- 设计上下文感知层,通过对比真实与伪造片段与全局上下文的距离来增强特征区分度。
- 在五个公开数据集上显著优于现有方法,实现更精确的时序伪造定位。
- 适合需要高精度局部篡改检测的多媒体取证场景,如视频真实性验证。
多媒体取证领域多数研究聚焦于深度伪造音视频内容的检测,取得了显著成果。然而,这些工作通常将检测视为分类任务,忽视了视频中仅部分片段被篡改的现实情况。小段伪造音视频嵌入真实视频中的时序伪造定位(TFL)仍具挑战性,更贴近实际应用场景。为此,本文提出一种通用上下文感知对比学习框架(UniCaCLF),利用监督对比学习通过异常检测识别并精确定位伪造时刻。创新性地引入上下文感知感知层,结合异构激活操作与自适应上下文更新机制,构建上下文感知对比目标,通过对比伪造与真实时刻到全局上下文的距离,提升伪造时刻特征的可区分性。进一步提出高效的上下文感知对比编码,以监督样本级方式强化真实与伪造时刻的特征差异,抑制跨样本干扰,从而提升时序伪造定位性能。在五个公开数据集上的大量实验表明,所提UniCaCLF显著优于当前最优算法。
原文摘要 · Abstract (English)
Most research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification task and ignore the case where partial segments of the video are tampered with. Temporal forgery localization (TFL) of small fake audio-visual clips embedded in real videos is still challenging and more in line with realistic application scenarios. To resolve this issue, we propose a universal context-aware contrastive learning framework (UniCaCLF) for TFL. Our approach leverages supervised contrastive learning to discover and identify forged instants by means of anomaly detection, allowing for the precise localization of temporal forged segments. To this end, we propose a novel context-aware perception layer that utilizes a heterogeneous activation operation and an adaptive context updater to construct a context-aware contrastive objective, which enhances the discriminability of forged instant features by contrasting them with genuine instant features in terms of their distances to the global context. An efficient context-aware contrastive coding is introduced to further push the limit of instant feature distinguishability between genuine and forged instants in a supervised sample-by-sample manner, suppressing the cross-sample influence to improve temporal forgery localization performance. Extensive experimental results over five public datasets demonstrate that our proposed UniCaCLF significantly outperforms the state-of-the-art competing algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。