arXiv:2605.00630cs.CVcs.MM2026-05被引 2

发现AI生成视频的跨模态时间指纹,提升检测泛化能力。

CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

论文配图:CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
图 1 · 摘自论文原文
  • 通过联合视觉文本嵌入与多粒度时序建模,捕捉跨模态对齐的时间特征。
  • 在4个数据集40个子集上达到新最好性能,跨生成器泛化能力强。
  • 适合关注AI视频真实性、跨模态分析的研究者和安全应用开发者。

先进AI视频合成技术的普及给数字视频真实性带来前所未有的挑战。现有AI生成视频(AIGV)检测方法主要关注单模态或时空伪影,忽略了视觉-文本跨模态空间中的丰富线索,尤其是语义对齐的时间稳定性。本文识别出AIGV中一种独特的指纹——跨模态时间伪影(CMTA)。真实视频因语义变化呈现自然的时间波动,而AIGV则因输入提示约束表现出不自然的稳定语义轨迹。为此,我们提出CMTA框架,通过联合跨模态嵌入与多粒度时序建模捕获此类时间特征。具体地,采用BLIP生成帧级图像描述,用CLIP提取对应视觉-文本表示;设计粗粒度分支利用GRU刻画跨模态对齐的时间波动;并行构建细粒度分支,通过Transformer编码器捕捉整合后的视觉-文本特征中的细微帧间变化。在包含GenVideo、EvalCrafter、VideoPhy、VidProM的四个大规模数据集共40个子集上的大量实验表明,本方法达到新最优性能,并展现出卓越的跨生成器泛化能力。代码与模型将开源于https://github.com/hwang-cs-ime/CMTA。

原文摘要 · Abstract (English)

The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video (AIGV) detection methods primarily focus on uni-modal or spatiotemporal artifacts, but they overlook the rich cues within the visual-textual cross-modal space, especially the temporal stability of semantic alignment. In this work, we identify a distinctive fingerprint in AIGVs, termed cross-modal temporal artifact (CMTA). Unlike real videos that exhibit natural temporal fluctuations in cross-modal alignment due to semantic variations, AIGVs display unnaturally stable semantic trajectories governed by given input prompts. To bridge this gap, we propose the CMTA framework, a cross-modal detection approach that captures these unique temporal artifacts through joint cross-modal embedding and multi-grained temporal modeling. Specifically, CMTA leverages BLIP to generate frame-level image captions and utilizes CLIP to extract corresponding visual-textual representations. A coarse-grained temporal modeling branch is then designed to characterize temporal fluctuations in cross-modal alignment with a GRU. In parallel, a fine-grained branch is constructed to capture intricate inter-frame variations from integrated visual-textual features with a Transformer encoder. Extensive experiments on 40 subsets across four large-scale datasets, including GenVideo, EvalCrafter, VideoPhy, and VidProM, validate that our approach sets a new state-of-the-art while exhibiting superior cross-generator generalization. Code and models of CMTA will be released at https://github.com/hwang-cs-ime/CMTA

AI视频检测跨模态分析生成内容验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。