细粒度评测发现大模型在生成视频检测中存在关键短板
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?
- 构建15维真实性维度的细粒度评测集,覆盖感知、光学、物理和时间线索
- 模型在感知线索上表现尚可,但在光学一致性和时序推理上严重不足
- 小模型有时比大模型更擅长特定检测任务,适合研究生成内容溯源
近年来,AI生成视频日益逼真复杂。大型视觉语言模型(LVLMs)虽在检测方面展现潜力,但现有评估多为二分类任务,依赖整体准确率等粗粒度指标,难以揭示模型实际优劣。为此,我们提出GenVideoLens,一个细粒度基准,支持对LVLM在生成视频检测中的能力进行多维度评估。该基准包含400个高度欺骗性的生成视频和100个真实视频,由专家标注了15个真实性维度,涵盖感知、光学、物理及时间线索。我们在该基准上评估了11个代表性LVLM。结果表明,模型在感知线索上表现较好,但在光学一致性、物理交互和时序因果推理上表现薄弱。不同模型在各维度表现差异显著,部分小型开源模型在特定线索上甚至优于大型专有模型。时序扰动实验显示,当前模型对时间信息利用有限。整体而言,GenVideoLens揭示了模型的关键能力缺口,为未来检测系统优化提供诊断依据。
原文摘要 · Abstract (English)
In recent years, AI-generated videos have become increasingly realistic and sophisticated. Meanwhile, Large Vision-Language Models (LVLMs) have shown strong potential for detecting such content. However, existing evaluation protocols largely treat the task as a binary classification problem and rely on coarse-grained metrics such as overall accuracy, providing limited insight into where LVLMs succeed or fail. To address this limitation, we introduce GenVideoLens, a fine-grained benchmark that enables dimension-wise evaluation of LVLM capabilities in AI-generated video detection. The benchmark contains 400 highly deceptive AI-generated videos and 100 real videos, annotated by experts across 15 authenticity dimensions covering perceptual, optical, physical, and temporal cues. We evaluate eleven representative LVLMs on this benchmark. Our analysis reveals a pronounced dimensional imbalance. While LVLMs perform relatively well on perceptual cues, they struggle with optical consistency, physical interactions, and temporal-causal reasoning. Model performance also varies substantially across dimensions, with smaller open-source models sometimes outperforming stronger proprietary models on specific authenticity cues. Temporal perturbation experiments further show that current LVLMs make limited use of temporal information. Overall, GenVideoLens provides diagnostic insights into LVLM behavior, revealing key capability gaps and offering guidance for improving future AI-generated video detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。