arXiv:2604.28078cs.CV2026-04被引 2

用专家评分体系提升视频美学质量,让生成视频更符合电影级审美。

AesRM: Improving Video Aesthetics with Expert-Level Feedback

论文配图:AesRM: Improving Video Aesthetics with Expert-Level Feedback
图 1 · 摘自论文原文
  • 构建三级美学框架,细化15个评判维度,实现系统化评估。
  • 建立包含2500对视频的专家标注数据集,支持多维度对比评测。
  • 提出可解释的推理模型AesRM-CoT,适合影视创作与内容优化场景。

尽管图像生成技术快速进步,但影视等真实应用仍需超越视觉保真度的美学表现,如色彩协调与电影级光影。现有研究多聚焦图像美学,常将美学简化为笼统概念。为此,本文提出分层评分体系,将视频美学分解为视觉美学(VA)、视觉保真度(VF)和视觉合理性(VP)三个核心维度,涵盖15项细粒度标准,如镜头构图。该框架支持构建大规模专家标注偏好数据集与评估基准AesVideo-Bench,包含约2500对视频及其在三维度上的专家标注。基于此,我们构建视频美学奖励模型家族AesRM:AesRM-Base直接预测成对偏好以提供高效后训练奖励;AesRM-CoT额外生成与15项标准对齐的思维链,增强评估可解释性。通过三阶段渐进训练:(1)基础美学能力学习,强化对中心构图等基本美学概念的识别;(2)冷启动阶段,对齐结构化推理协议;(3)GRPO进一步提升评估准确性。为提升AesRM-CoT性能,我们引入自一致性思维链合成方法,并设计基于思维链的过程奖励。大量实验表明,AesRM在多个美学基准上优于基线,且具备更低的位置偏差。最后,将Wan2.2与AesRM对齐后,显著提升美学表现。

原文摘要 · Abstract (English)

Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors and cinematic lighting, beyond visual fidelity. Prior work on visual aesthetics largely focuses on images, often reducing aesthetics to coarse definitions, e.g., visual pleasure, without a rigorous and systematic evaluation. To improve video aesthetics, we propose a hierarchical rubric that decomposes video aesthetics into three core dimensions, Visual Aesthetics (VA), Visual Fidelity (VF), and Visual Plausibility (VP), with 15 fine-grained criteria, e.g., shot composition. This framework enables a large-scale expert-annotated preference dataset and an evaluation benchmark, AesVideo-Bench, containing about 2500 video pairs with expert annotations on VA, VF, and VP. We then build a family of Video Aesthetic Reward Models (AesRM): AesRM-Base, which directly predicts pairwise preferences on these dimensions to provide efficient post-training rewards, and AesRM-CoT, which additionally generates CoT aligned with all 15 criteria to improve assessment interpretability. Specifically, we train AesRM with a three-stage progressive scheme: (1) Atomic Aesthetic Capability Learning, which strengthens AesRM's recognition of fundamental aesthetic concepts, e.g., accurately identifying centered composition; (2) Cold-Start, aligning the model with structured reasoning protocols; and (3) GRPO, further improving evaluation accuracy. To enhance AesRM-CoT, we additionally propose self-consistency-based CoT synthesis to improve CoT quality and design CoT-based process rewards during GRPO. Extensive experiments show AesRM outperforms baselines on multiple aesthetics benchmarks and is more robust, with lower position bias. Finally, we align Wan2.2 with AesRM and observe clear aesthetic gains over existing aesthetic reward models.

视频美学奖励模型专家标注可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。