arXiv:2605.17311cs.CV2026-05被引 1

融合频谱与语义特征,提升高保真合成视频检测准确率

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection

论文配图:SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection
图 1 · 摘自论文原文
  • 设计频谱模块提取傅里叶变换后的高频特征
  • 在自研基准上达95.59%准确率,优于现有方法
  • 适合关注生成视频安全的AI安全研究者

近期商业视频生成模型(如Sora和Veo)展现出极高的视觉保真度,使得可靠的AI生成视频检测愈发关键,以防止合成内容被误认为真实视频并用于误导信息传播。然而,现有检测器常因过度依赖日益逼真的语义特征,忽视细微的频谱伪影而失效。本文提出SpecSem-Net,首个专为高保真生成视频设计的语义引导频谱去噪框架。具体地,我们构建频谱模块,通过基于傅里叶变换的滤波提取高频特征;为减少频谱噪声导致的误判,引入门控融合机制,自适应融合语义上下文,有效抑制噪声。此外,为评估检测器在顶级生成模型上的表现,我们构建涵盖5个顶尖商用生成器的综合性基准。大量实验表明,SpecSem-Net在自建基准和公开数据集上分别达到87.25%和95.59%的准确率,显著优于现有方法。

原文摘要 · Abstract (English)

The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasingly essential to prevent synthetic content from being indistinguishable from real videos and exploited for disinformation. However, existing detectors often fail due to an over-reliance on increasingly realistic semantic features, neglecting subtle spectral artifacts. In this paper, we propose SpecSem-Net, the first framework to introduce a semantic-guided spectral denoising mechanism specifically for high-fidelity AI-generated video detection. Specifically, we design a spectral module to extract high-frequency features via Fourier-Transform based filtering. Furthermore, to reduce misjudgments arising from spectral noise, we employ a Gated Merging Mechanism to adaptively fuse semantic context, effectively mitigating spectral noise. Additionally, to evaluate detector performance on the latest top-tier generative models, we construct a comprehensive benchmark comprising 5 SOTA commercial generators. Extensive experiments demonstrate that SpecSem-Net outperforms existing methods, achieving accuracies of 87.25% and 95.59% on our benchmark and public datasets, respectively.

视频生成检测频谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。