测试1.8万段生成视频,发现现有检测器在真实危机场景下普遍失效。
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

- 用真实视频作锚点构建1.8万段生成视频数据集
- 检测器对高质量生成视频的识别率不足50%且不稳定
- 社交传播使检测更难,人和机器都容易被误导
近期视频生成技术可逼真再现战争、灾难、公共危机等现实事件,带来严重虚假信息风险。现有基准对检测器与生成器在真实场景下的表现评估有限,包括检测难度随生成条件变化情况、人类对生成视频的感知差异,以及社交传播中检测可靠性问题。为此,我们提出RA-Bench,一个以真实视频为锚点的AI生成视频检测基准。该数据集包含17,886段视频:1,830个真实视频锚点(涵盖10类社会风险事件)及16,056段来自4个开源和5个闭源生成器的合成视频。基于此,我们从三个维度展开评估:首先,在七种传统检测器、十种零样本多模态模型及两种针对生成视频检测微调的MLLMs上测试泛化能力,结果表明三类检测器均无法在所有实例上保持一致性能;其次,分析生成质量、条件信息和采样种子对检测性的影响,发现不同生成特性对各类检测器影响各异,但源级检测模式在不同种子间保持稳定;最后,研究人类真实性判断与社交传播中的检测可靠性,发现误导性强的视频同样难以被当前检测器识别,且社交传播进一步降低检测效果。综合表明,现有方法在应对真实世界高保真生成视频时仍存在显著局限,亟需更具鲁棒性的检测技术。
原文摘要 · Abstract (English)
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。