arXiv:2608.21839cs.CV2026-08

用检查清单提升文本生成视频的评分可靠性

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

论文配图:FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
图 1 · 摘自论文原文
  • 先逐项检查视频内容再打分,确保每条评价有视觉证据支持
  • 构建9万条标注数据,实现比现有模型更低的评分误差
  • 适合需要精准评估视频质量的研究者与开发者

可靠的奖励模型对文本到视频的评估与对齐至关重要。然而,评估精度与推理效率之间的权衡对训练监督质量提出了高要求。现有方法常依赖固定标准的全局评判或开放式推理,导致检查不全、理由失真、归因混乱。我们提出FIRM-Video,一种基于‘检查后再打分’原则的统一清单驱动数据构建框架:为每个维度设计特定检查清单,逐项验证其在时间序列视觉证据中的成立性,仅聚合已验证的判断结果。针对指令遵循,将提示分解为加权原子需求;针对世界一致性,构建基于可见实体与动作的目标特异性检查;针对感知质量,采用通用视觉缺陷分类体系。经验证的标准与评分被转化为自然语言分析,用于端到端奖励建模。我们构建了包含88,044个维度特异性实例的FIRM-Video-90K数据集(来自29,348段视频),并推出包含750条点对点人工标注的FIRM-Video-Bench(覆盖250段视频)。基于Qwen3-VL的FIRM-Video-8B在FIRM-Video-Bench上取得最优整体MAE,且在三种视频生成器的Best-of-8采样中,持续获得最高VBench总分、质量和语义得分。

原文摘要 · Abstract (English)

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

视频评估奖励模型数据构建多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。