arXiv:2506.18564cs.CV2025-06AAAI被引 34

用渐进式强化学习教视觉语言模型理解AI生成视频质量

VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning

  • 分三阶段训练:先学图像质量,再学时序特征,最后与生成模型联合优化
  • 多维度奖励机制使评分在偏好对比、多维打分和自然视频评估中均领先
  • 适合需要高质量视频生成反馈的AI研发者,尤其关注生成内容评估

近年来,文本到视频生成模型取得了显著进展。然而,由于泛化能力有限、缺乏时序感知、严重依赖大规模标注数据,且难以与生成模型有效交互,AI生成视频的质量评估仍具挑战。现有方法多采用监督微调视觉语言模型(VLMs),往往需大量标注数据,且理解与生成过程脱节。为此,我们提出VQ-Insight,一种新型推理式VLM框架,用于AIGC视频质量评估。其核心包括:(1) 渐进式视频质量学习策略,融合图像质量预热、任务特定时序学习及与视频生成模型的联合优化;(2) 设计多维度评分奖励、偏好对比奖励与时序建模奖励,提升评估的泛化性与专业性。大量实验表明,VQ-Insight在偏好对比、多维度评分和自然视频评分任务中持续优于当前最优基线,显著提升视频生成质量。

原文摘要 · Abstract (English)

Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-generated videos remains challenging due to limited generalization, lack of temporal awareness, heavy reliance on large-scale annotated datasets, and the lack of effective interaction with generation models. Most current approaches rely on supervised finetuning of vision-language models (VLMs), which often require large-scale annotated datasets and tend to decouple understanding and generation. To address these shortcomings, we propose VQ-Insight, a novel reasoning-style VLM framework for AIGC video quality assessment. Our approach features: (1) a progressive video quality learning scheme that combines image quality warm-up, general task-specific temporal learning, and joint optimization with the video generation model; (2) the design of multi-dimension scoring rewards, preference comparison rewards, and temporal modeling rewards to enhance both generalization and specialization in video quality evaluation. Extensive experiments demonstrate that VQ-Insight consistently outperforms state-of-the-art baselines in preference comparison, multi-dimension scoring, and natural video scoring, bringing significant improvements for video generation tasks.

视频质量评估视觉语言模型生成模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。