通过反演扩散模型初始噪声,区分真实与生成视频。
DBINDS -- Can Initial Noise from Diffusion Model Inversion Help Reveal AI-Generated Videos?
- 利用扩散模型反演提取初始噪声序列差异。
- 在单生成器训练下跨多生成器检测准确率达87.3%。
- 适合内容安全、伪造视频检测场景使用。
AI生成视频快速发展,对内容安全和取证分析构成严峻挑战。现有检测器主要依赖像素级视觉线索,在未见生成器上泛化能力差。本文提出DBINDS,一种基于扩散模型反演的检测方法,分析潜在空间动态而非像素。我们发现,通过扩散模型反演恢复的初始噪声序列在真实与生成视频间存在系统性差异。基于此,DBINDS构建初始噪声差异序列(INDS),提取多域、多尺度特征。结合特征优化与贝叶斯调优的LightGBM分类器,DBINDS仅在单一生成器上训练,即在GenVidBench上实现强跨生成器性能,展现出良好的泛化能力与小样本鲁棒性。
原文摘要 · Abstract (English)
AI-generated video has advanced rapidly and poses serious challenges to content security and forensic analysis. Existing detectors rely mainly on pixel-level visual cues and generalize poorly to unseen generators. We propose DBINDS, a diffusion-model-inversion based detector that analyzes latent-space dynamics rather than pixels. We find that initial noise sequences recovered by diffusion inversion differ systematically between real and generated videos. Building on this, DBINDS forms an Initial Noise Difference Sequence (INDS) and extracts multi-domain, multi-scale features. With feature optimization and a LightGBM classifier tuned by Bayesian search, DBINDS (trained on a single generator) achieves strong cross-generator performance on GenVidBench, demonstrating good generalization and robustness in limited-data settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。