构建首个可定位真假痕迹的视频假象评测集,让AI理解人类如何识破深度伪造。
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
- 构建多模态大模型,学习人类对视频假迹的感知与定位。
- 4.3万条标注覆盖3300个高质量生成视频,含时空定位与自然语言解释。
- 模型在识别、定位、解释三方面均显著超越GPT-5,适合可信生成研究。
尽管视频生成模型发展迅速,但一个关键问题长期被忽视:人类能否识别生成视频中的真实痕迹(即时空上可定位的机器生成痕迹)?我们提出DeeptraceReward,首个细粒度、时空感知的基准数据集,用于标注人类对生成视频中假象痕迹的感知。该数据集包含3300个高质量生成视频的4.3万条详细标注,每条标注均提供自然语言解释,精确定位含假迹的边界框区域,并标记起止时间。我们归纳出9类导致人类识别为伪造的主要痕迹类型,并训练多模态语言模型作为奖励模型以模仿人类判断与定位。在DeeptraceReward上,我们的7B规模奖励模型在假迹识别、定位和解释三项任务上的平均性能比GPT-5高出34.7%。值得注意的是,人类判断存在明显难度梯度:二分类(真/假)远易於细粒度痕迹检测;而痕迹检测中,自然语言解释最易,空间定位次之,时间标注最难。DeeptraceReward为社会敏感且可信的视频生成提供了严格测试平台与训练信号。
原文摘要 · Abstract (English)
Can humans identify AI-generated (fake) videos and provide grounded reasons? While video generation models have advanced rapidly, a critical dimension -- whether humans can detect deepfake traces within a generated video, i.e., spatiotemporal grounded visual artifacts that reveal a video as machine generated -- has been largely overlooked. We introduce DeeptraceReward, the first fine-grained, spatially- and temporally- aware benchmark that annotates human-perceived fake traces for video generation reward. The dataset comprises 4.3K detailed annotations across 3.3K high-quality generated videos. Each annotation provides a natural-language explanation, pinpoints a bounding-box region containing the perceived trace, and marks precise onset and offset timestamps. We consolidate these annotations into 9 major categories of deepfake traces that lead humans to identify a video as AI-generated, and train multimodal language models (LMs) as reward models to mimic human judgments and localizations. On DeeptraceReward, our 7B reward model outperforms GPT-5 by 34.7% on average across fake clue identification, grounding, and explanation. Interestingly, we observe a consistent difficulty gradient: binary fake v.s. real classification is substantially easier than fine-grained deepfake trace detection; within the latter, performance degrades from natural language explanations (easiest), to spatial grounding, to temporal labeling (hardest). By foregrounding human-perceived deepfake traces, DeeptraceReward provides a rigorous testbed and training signal for socially aware and trustworthy video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。