用光学计算并行检测多路深伪视频,速度快能耗低且抗干扰。
Scalable, Energy-Efficient Optical-Neural Architecture for Multiplexed Deepfake Video Detection
- 数字前端+光学后端混合架构,通过光调制器并行处理15路视频
- 单次光传播中完成15个视频检测,准确率达97.79%,灵敏度99.86%
- 适合高并发、低功耗的实时深伪内容检测场景
AI生成视觉内容的快速泛滥催生了高效可信的深伪检测系统需求。现有深度学习方法依赖高算力与高能耗的推理算法,限制了可扩展性。本文提出一种混合数字-模拟深伪视频检测框架,结合轻量级数字前端与空间复用的光学解码后端,利用可编程空间光调制器实现大规模并行模拟推理。通过单次光传播过程同时处理15个以上视频流,系统在保持高吞吐量的同时显著降低计算成本。我们在涵盖人脸替换、真实世界深伪视频及全生成视频的不同数据集上验证了该系统。在可见光波段的空间复用实验设置下,于Celeb-DF数据集上以单次光路并行测试15个视频,实现了97.79%的平均检测准确率、99.86%的敏感度和95.72%的特异性。该复用光学解码器还对多种视频退化、噪声、压缩、对齐偏差及黑盒对抗攻击表现出强鲁棒性。结果表明,将光学计算融入AI推理可同时提升吞吐量、能效与对抗鲁棒性——这三项优势在纯数字系统中难以兼得。
原文摘要 · Abstract (English)
The rapid proliferation of AI-generated visual media has created an urgent need for efficient, trustworthy deepfake detection systems. However, existing deep learning-based detection methods rely on computationally intensive and energy-demanding inference algorithms, limiting their scalability. Here, we present a hybrid digital-analog deepfake video detection framework that combines a lightweight digital front-end with a spatially multiplexed optical decoding back-end for massively parallel analog inference through a programmable spatial light modulator. By simultaneously processing 15 or more video streams within a single optical propagation pass, the system enables high-throughput and accurate video-level authenticity prediction at reduced computational cost compared with purely digital methods. We validated this hybrid deepfake video processor using different datasets spanning classical face-swapping, real-world deepfake recordings, and fully AI-generated videos. Using a spatially multiplexed experimental set-up operating in the visible spectrum, we achieved average deepfake detection accuracy, sensitivity and specificity of 97.79%, 99.86% and 95.72%, respectively, on the Celeb-DF video dataset with 15 videos tested in parallel in a single optical pass per inference. The multiplexed optical decoder also demonstrates resilience against various types of video degradation, noise, compression, experimental misalignments and black-box adversarial attacks. Our results show that integrating optical computation into AI inference enables simultaneous gains in throughput, energy efficiency, and adversarial robustness - three properties that are difficult to achieve together in purely digital systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。