arXiv:2602.08828cs.CV2026-02被引 12

用感知预训练强化学习,提升对生成视频的检测能力

VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

  • 通过时空定位和自监督计数任务,增强模型感知能力
  • 在9个生成器的3000段视频上达到更均衡的检测效果
  • 适合关注生成内容安全与多模态检测的研究者

视频生成技术的快速发展带来了日益严峻的安全风险,可靠检测愈发重要。本文提出VideoVeritas框架,融合细粒度感知与事实推理能力。观察到当前多模态大模型虽有强推理能力,但感知精度有限。为此,我们引入联合偏好对齐与感知预训练强化学习(PPRL):不直接优化检测任务,而是在强化学习阶段采用通用时空定位和自监督物体计数作为预训练任务,显著提升检测性能。为支持稳健评估,我们构建了MintVid数据集,包含来自9个顶尖生成器的3000段视频,以及一个含有内容事实错误的真实世界子集。实验表明,现有方法普遍偏向表面推理或机械分析,而VideoVeritas在多个基准测试中实现了更均衡的表现。

原文摘要 · Abstract (English)

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large language models (MLLMs) exhibit strong reasoning capacity, their granular perception ability remains limited. To mitigate this, we introduce Joint Preference Alignment and Perception Pretext Reinforcement Learning (PPRL). Specifically, rather than directly optimizing for detection task, we adopt general spatiotemporal grounding and self-supervised object counting in the RL stage, enhancing detection performance with simple perception pretext tasks. To facilitate robust evaluation, we further introduce MintVid, a light yet high-quality dataset containing 3K videos from 9 state-of-the-art generators, along with a real-world collected subset that has factual errors in content. Experimental results demonstrate that existing methods tend to bias towards either superficial reasoning or mechanical analysis, while VideoVeritas achieves more balanced performance across diverse benchmarks.

视频检测生成安全强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。