arXiv:2602.05202cs.CV2026-02被引 1

用生成模型当评分器,更准地判断视频质量。

GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

  • 把视频生成模型改造成能量模型,靠对比学习判断视频好坏
  • 仅用3万标注数据,性能超越现有方法6到65倍
  • 用扰动生成假视频,逼模型学真实时空特征

对齐视频生成模型与人类偏好仍具挑战:现有方法依赖视觉语言模型(VLM)进行奖励建模,但难以捕捉细微的时间动态。本文提出新思路:将本身擅长建模时间结构的视频生成模型重新用作奖励模型。我们提出生成式变压器自监督视频判别器(GT-SVJ),将先进视频生成模型转化为强时序感知的奖励模型。核心思想是将生成模型重构为能量模型(EBM),通过对比学习使高质量视频获得低能量、劣质视频获得高能量,从而精准区分视频质量。为防止模型依赖真实与生成视频间的表面差异,我们设计了包含时序切片、特征交换和帧打乱的合成负样本,模拟真实但细微的视觉退化,迫使模型学习有意义的时空特征。GT-SVJ在GenAI-Bench和MonteBench上达到当前最佳性能,仅需30K人工标注,比现有VLM方法减少6至65倍。

原文摘要 · Abstract (English)

Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle to capture subtle temporal dynamics. We propose a fundamentally different approach: repurposing video generative models, which are inherently designed to model temporal structure, as reward models. We present the Generative-Transformer-based Self-Supervised Video Judge (\modelname), a novel evaluation model that transforms state-of-the-art video generation models into powerful temporally-aware reward models. Our key insight is that generative models can be reformulated as energy-based models (EBMs) that assign low energy to high-quality videos and high energy to degraded ones, enabling them to discriminate video quality with remarkable precision when trained via contrastive objectives. To prevent the model from exploiting superficial differences between real and generated videos, we design challenging synthetic negative videos through controlled latent-space perturbations: temporal slicing, feature swapping, and frame shuffling, which simulate realistic but subtle visual degradations. This forces the model to learn meaningful spatiotemporal features rather than trivial artifacts. \modelname achieves state-of-the-art performance on GenAI-Bench and MonteBench using only 30K human-annotations: $6\times$ to $65\times$ fewer than existing VLM-based approaches.

视频生成自监督奖励建模时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。