通过分层边界建模定位音视频深度伪造片段,精准捕捉细微篡改。
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
- 构建分层网络,融合音视频特征并建模时间边界
- 在小范围篡改场景下实现高精度定位,提升召回率与精确率
- 适合需细粒度检测伪造内容的研究者与安全团队
在内容驱动的部分篡改场景下,音视频时序深度伪造定位仍具挑战性。此时伪造区域通常仅持续数帧,其余部分与原始内容一致。为此,本文提出分层边界建模网络(HBMNet),包含三个模块:音视频特征编码器提取判别性帧级表征,粗略候选边界生成器预测可能的篡改区域,细粒度概率生成器利用双向边界-内容概率对候选区域进行精修。从模态角度,通过专用编码与融合,并结合帧级监督以增强判别能力;从时间角度,整合多尺度线索与双向边界-内容关系。实验表明,编码与融合主要提升精确率,帧级监督则改善召回率。各模块(音视频融合、时间尺度、双向性)贡献互补,共同提升定位性能。HBMNet优于BA-TFD与UMMAFormer,且在更多训练数据下展现出更强可扩展性。
原文摘要 · Abstract (English)
Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。