arXiv:2605.07232cs.CV2026-05

多模态联合建模,精准定位AI生成视频的局部篡改

Towards multi-modal forgery representation learning for AI-generated video detection and localization

论文配图:Towards multi-modal forgery representation learning for AI-generated video detection and localization
图 1 · 摘自论文原文
  • 融合视觉、音频与语义信息的多模态架构
  • 在多个数据集上实现更高检测准确率与精细时间定位
  • 适合需要高精度伪造内容定位的安防与内容审核场景

生成式AI的快速发展使大规模视频创作变得普及。部分篡改的跨模态(视觉与音频)AI生成视频正带来日益严重的语义失真与滥用风险,亟需可靠的检测工具。现有方法多受限于单一或局部模态建模,且缺乏细粒度的时间伪造定位能力。为此,本文提出一种核心架构,联合集成语言-多模态(LMM)语义分支、时空(ST)视觉分支与多尺度局部伪造(PS)音频分支。该多模态方法可同时实现对部分篡改的AI生成视频伪造内容的检测与细粒度时间定位。大量实验表明,该方法优于当前最先进的检测技术。

原文摘要 · Abstract (English)

Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and audio channels, pose escalating risks of semantic distortion and misuse, which motivates the need for reliable detection tools. Most existing AI-generated video detectors remain limited by single- or partial-modality of data modeling and the lack of fine-grained temporal forgery localization. To address these challenges, our primary novelty introduces a core architecture that jointly integrates an LMM semantic branch with a spatio-temporal (ST) visual branch and a multi-scale partial-spoof (PS) audio branch. This multi-modal approach enables simultaneous detection and fine-grained temporal localization of partially manipulated AI-generated video forgeries. Extensive experiments show that this approach outperforms existing state-of-the-art methods.

AI伪造检测多模态学习视频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。