arXiv:2605.23271cs.CVcs.AI2026-05被引 4

为专业影视视频生成设计了更懂美学的评估体系

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

论文配图:EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
图 1 · 摘自论文原文
  • 按电影制作流程构建评估分类体系
  • 用专家标注数据训练出能推理的视觉语言模型
  • 适合研究高质量视频生成与智能评估的人

生成式视频基础模型快速发展,推动领域迈向专业级影视合成。为实现高要求质量,社区转向强化学习与智能体工作流。然而,可靠评估已成为关键瓶颈。现有基准多仅衡量“是否正确”(基本指令遵循),却忽视“是否优质”(影像质量、表演与美学)。当前自动指标缺乏领域严谨性,导致机器评分与人类审美感知间存在严重可信度差距。为此,我们提出EvalVerse——一个全流程感知、专家校准的综合评估框架。首先,将领域知识组织为契合专业电影制作流程(前期、拍摄、后期)的评估分类体系;其次,通过大规模人工标注构建专家判断数据集;第三,采用专家校准微调策略,将知识注入视觉语言模型(VLM),使其具备显式链式思考能力。相比以往方法,EvalVerse不仅兼容基础“正确性”指标,更显著扩展至“优质性”维度,并覆盖复杂多镜头序列与音画融合任务。由此提供细粒度诊断信号,超越静态排行榜,为未来奖励模型与评估智能体奠定基础。

原文摘要 · Abstract (English)

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic workflows. However, reliable evaluation has emerged as a critical bottleneck. Existing benchmarks predominantly evaluate ''whether it is right'' (basic prompt-following) while fundamentally neglecting ''whether it is good'' (cinematic quality, acting, and aesthetics). Furthermore, current automated metrics lack the domain-specific rigor required to provide trustworthy signals, creating a severe credibility gap between human aesthetic perception and machine scoring. To bridge this gap, we introduce EvalVerse, a comprehensive, pipeline-aware, and expert-calibrated evaluation framework. We treat video generation assessment not merely as an engineering task, but as a core scientific problem: the systematic digitization of subjective cinematic expertise. First, we organize domain knowledge into an evaluation taxonomy aligned with the professional filmmaking workflow (pre-production, production, and post-production). Second, we distill human expert judgments into a curated dataset with large-scale human annotations. Third, we inject this knowledge into Vision-Language Models (VLMs) through an expert-calibrated fine-tuning strategy, enabling the VLM to perform explicit Chain-of-Thought reasoning. Compared to previous works, EvalVerse not only retains compatibility with foundational ''rightness'' metrics, but also significantly expands the criteria to ''goodness'' and broaden the task coverage to complex multi-shot sequencing and audio-visual integration. Consequently, by providing granular diagnostic signals, EvalVerse transcends a static leaderboard and establishes a fundamental infrastructure for future work, such as reward models and evaluator agent.

视频生成评估框架影视合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。