通过音视频协同与边界校准,精准定位长视频中的伪造片段。
EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

- 分阶段音视频协同学习,捕捉稀疏伪造的深层语义痕迹。
- 引入边界感知精修策略,提升伪造段落边界的定位精度。
- 轻量化设计适合实际部署,适用于真实场景的伪造检测。
人工智能生成内容的快速扩散亟需可靠的多模态取证技术。除了视频级二分类,精确识别长视频中稀疏分布的伪造片段仍是关键挑战,尤其当篡改痕迹细微、跨模态信号弱且时间上分散时。为此,我们提出EVAS,一种端到端的多模态时序伪造定位框架。核心采用多阶段音视频协同机制,促进渐进式跨模态交互,学习深层多模态取证表征并捕捉稀疏篡改的高阶语义痕迹。同时引入边界感知精修策略,通过无效帧掩码抑制模糊区域,强化过渡区域预测。采用解耦训练范式搭配辅助头,分离表征学习与推理目标,提升模型泛化性与稳定性。此外,引入轻量级HourglassFFN以降低计算开销。大量实验表明,EVAS在三个基准数据集上均达到最先进的平均定位准确率和平均召回率,验证了其在细粒度时序伪造定位中的有效性。
原文摘要 · Abstract (English)
The rapid proliferation of artificial intelligence-generated content necessitates reliable multimodal forensics. Beyond video-level binary classification, precisely localizing sparsely distributed forged segments in long-form videos remains a critical challenge. This task is particularly difficult when manipulations are subtly embedded and cross-modal signals are weak and temporally diffuse. To address these challenges, we propose EVAS, an end-to-end multimodal framework for temporal forgery localization. At its core, a Multi-Stage Audio-Visual Synergy mechanism facilitates progressive cross-modal interaction to learn deep multimodal forensic representations and capture high-order semantic traces of sparse manipulations. Furthermore, we introduce a Boundary-Aware Refinement strategy to achieve steered boundary calibration. By incorporating invalid-frame masking, this strategy suppresses ambiguous regions and sharpens transition predictions. We adopt a decoupled training paradigm with auxiliary heads to disentangle representation learning from inference objectives, enhancing model generalization and stability. Additionally, a lightweight HourglassFFN is incorporated to reduce computational overhead. Extensive experiments demonstrate that EVAS achieves state-of-the-art average localization accuracy and average recall across three benchmark datasets, validating its effectiveness for fine-grained temporal forgery localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。