提出六控制审计协议与工具包,让AI视频检测评估更可信。
Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit

- 设计六项标准控制,杜绝评估中的干扰因素。
- 多方法在严格审计后性能大幅下降,部分仅剩个位数召回率。
- 开源工具包整合14个检测器,支持可复现的公平对比。
当前AI生成视频检测基准(如GenVidBench、AIGVDBench)虽为行业标准,但多数评估未控制干扰变量,导致泛化性能被夸大。实验表明,一个仅依赖片段长度的简单分类器在未经审计的测试中可达0.998的LOGO AUC,却完全忽略运动信息。20篇论文综述发现无一使用全部六项控制,因此本文提出统一审计协议,并应用于六个代表性特征源(三个已发表检测器和三个重用信号源),在AIGVDBench上进行跨数据集重测。审计结果揭示:简单分类器性能降至近随机水平(0.529),CLIP基线暴露数据集身份泄露问题;而2025年新检测器WaveRep在分布外检测中仍保持0.996的LOGO AUC,且真实样本间一致性接近随机。在部署级假阳性率0.1%下,多个高AUC方法召回率跌至个位数,排行榜顺序改变。建议采用(AUC、超出底线的裕度、工作点召回率、校准性)组合指标替代单一数值。引入时间谱(TemporalSpec)作为白盒正向控制,通过跨底座特征融合(XSFF)实现真正互补性。发布VidAudit工具包,集成14个检测器、统一API、排行榜及元数据,可在https://github.com/KurbanIntelligenceLab/vidaudit获取。该协议与工具推动评估从排名导向转向对测量有效性的验证。
原文摘要 · Abstract (English)
AI-generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontrolled confounds that can inflate reported generalization. As an existence proof, a three-feature clip-length classifier reaches a leave-one-generator-out (LOGO) AUC of 0.998 on GenVidBench under unaudited evaluation, while measuring nothing about motion. A 20-paper survey finds none applying all six standard controls that would catch this, so we combine them into an audited protocol and apply it to six representative feature sources (three published detectors and three repurposed signal sources), re-running it cross-dataset on AIGVDBench. The audit both debunks and certifies: the trivial classifier collapses to near chance (0.529), a CLIP baseline is caught carrying dataset identity, and the 2025 forensic detector WaveRep clears the floor at out-of-distribution LOGO AUC 0.996 with chance-level real-vs-real coherence. At a deployable FPR of 0.1%, multiple high-AUC methods fall to single-digit recall and the leaderboard order changes, so we recommend an audited tuple (AUC, above-floor margin, operating-point recall, and calibration) over a single number. As a white-box positive control, we add TemporalSpec (codec motion vectors); via cross-substrate feature fusion (XSFF), a second substrate adds genuine complementarity that survives the audit. We release VidAudit, to our knowledge the largest unified and audited detector collection for this task, providing 14 detectors behind one plugin API, a leaderboard, and Croissant metadata, available at https://github.com/KurbanIntelligenceLab/vidaudit. Together, the protocol and toolkit move evaluation from leaderboard rank toward whether a result measures what it claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。