融合多维度线索,精准定位AI伪造视频的细微痕迹
DYMAPIA: A Multi-Domain Framework for Detecting AI-based Video Manipulation

- 结合频域、纹理与运动一致性信息生成动态异常掩码
- 在多个数据集上准确率和F1值超99%,支持实时检测
- 适合需要快速验证真伪的媒体审核与反虚假信息场景
AI生成内容快速发展,引发对内容真实性和数字信任的担忧。我们提出DYMAPIA,一种多领域深度伪造检测框架,通过融合空间、频谱和时间线索,捕捉视觉数据中细微的篡改痕迹。系统利用傅里叶频谱、局部纹理描述符、边缘不规则性及光流一致性生成动态异常掩码,精确定位被篡改区域。这些掩码指导一个轻量级分类器DistXCNet,该模型基于Xception架构并采用深度可分离卷积优化,实现快速、区域聚焦的分类。联合设计在FF++、Celeb-DF和VDFD基准上均取得超过99%的准确率和F1分数,同时保持模型紧凑,适用于实时应用。相较于现有全帧和多领域检测器,DYMAPIA展现出出色的性能与部署潜力,可直接用于媒体真实性验证、虚假信息防御和安全内容过滤等时效性任务。
原文摘要 · Abstract (English)
AI-generated media are advancing rapidly, raising pressing concerns for content authenticity and digital trust. We introduce DYMAPIA, a multi-domain Deepfake detection framework that fuses spatial, spectral, and temporal cues to capture subtle traces of manipulation in visual data. The system builds dynamic anomaly masks by combining evidence from Fourier spectra, local texture descriptors, edge irregularities, and optical flow consistency, which highlight tampered regions with fine spatial accuracy. These masks guide DistXCNet, a lightweight classifier distilled from Xception and optimized with depthwise separable convolutions for fast, region-focused classification. This joint design achieves state-of-the-art results, with accuracy and F1-scores exceeding 99\% on FF++, Celeb-DF, and VDFD benchmarks, while keeping the model compact enough for real-time use. Beyond outperforming existing full-frame and multidomain detectors, DYMAPIA demonstrates deployment readiness for time-critical forensic tasks, including media verification, misinformation defense, and secure content filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。