arXiv:2607.22632cs.AI2026-07

构建多维度视频编辑评估模型,提升自动化剪辑优化效果。

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

论文配图:VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
图 1 · 摘自论文原文
  • 定义六维评价体系,融合专业创作者意见
  • 基于10万条剪辑数据训练,性能超越GPT-5等模型
  • 支持细粒度打分与可操作反馈,适合内容创作者

短视频作为个性化叙事媒介迅速兴起,亟需自动化系统评估与优化剪辑方案。然而,视频评估高度主观,缺乏统一标准、数据集与基准测试。为此,我们联合专业创作者与产品经理,构建涵盖创意、一致性、概念设计、摄影、叙述与节奏共六维度的综合评价框架。随后,我们收集了包含10万条剪辑样本的大规模数据集,并建立专用基准VRMBench,用于评估多模态大语言模型(MLLMs)的视频奖励能力。在此基础上,提出VlogReward模型,能提供细粒度多维度评分及可执行反馈,支持迭代优化。技术上,通过引入可调组间对比奖励,改进组相对策略优化(GRPO)框架,缓解标准GRPO的“方向盲区”问题,使模型更准确区分不同质量剪辑。实验表明,VlogReward在多项指标上显著优于现有MLLMs,包括GPT-5和Gemini-3-Pro。本研究旨在助力创作者并推动自动化视频评估与优化系统的建设。

原文摘要 · Abstract (English)

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.

视频评估多模态生成优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。