arXiv:2601.04033cs.CV2026-01被引 4

专门评估生成视频结构畸变,提升对异常物体与互动的检测能力。

Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model

  • 基于帧级推理,通过分步思考识别视频中的结构畸变。
  • 在自建数据集上训练,量化评估畸变并提供可解释的归因标签。
  • 适合关注生成视频质量、尤其是结构合理性的研究者使用。

近期视频奖励模型和后训练策略提升了文本到视频(T2V)生成质量,但通常只关注视觉质量、运动质量和文本对齐,忽视了异常物体外观与交互等关键结构性畸变,影响整体生成视频质量。为此,我们提出REACT,一种专用于生成视频结构性畸变评估的帧级奖励模型。REACT通过推理视频帧,分配点级分数与归因标签,聚焦于畸变识别。我们构建了一个大规模人工偏好数据集,基于新提出的结构性畸变分类体系标注,并利用高效的思维链(CoT)合成管道生成额外数据。REACT采用两阶段训练框架:(1)使用掩码损失进行监督微调以注入领域知识;(2)采用组相对策略优化(GRPO)与成对奖励进行强化学习,提升推理能力并使输出分数与人类偏好对齐。推理时引入动态采样机制,聚焦最可能含畸变的帧。我们还提出了REACT-Bench,一个生成视频畸变评估基准。实验表明,REACT能有效补充现有模型,在结构性畸变评估中实现准确量化与可解释归因分析。

原文摘要 · Abstract (English)

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural distortions, such as abnormal object appearances and interactions, which can degrade the overall quality of the generative video. To address this gap, we introduce REACT, a frame-level reward model designed specifically for structural distortions evaluation in generative videos. REACT assigns point-wise scores and attribution labels by reasoning over video frames, focusing on recognizing distortions. To support this, we construct a large-scale human preference dataset, annotated based on our proposed taxonomy of structural distortions, and generate additional data using a efficient Chain-of-Thought (CoT) synthesis pipeline. REACT is trained with a two-stage framework: (1) supervised fine-tuning with masked loss for domain knowledge injection, followed by (2) reinforcement learning with Group Relative Policy Optimization (GRPO) and pairwise rewards to enhance reasoning capability and align output scores with human preferences. During inference, a dynamic sampling mechanism is introduced to focus on frames most likely to exhibit distortion. We also present REACT-Bench, a benchmark for generative video distortion evaluation. Experimental results demonstrate that REACT complements existing reward models in assessing structutal distortion, achieving both accurate quantitative evaluations and interpretable attribution analysis.

视频生成奖励模型畸变检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。