arXiv:2602.21835cs.CV2026-02被引 10

首个统一评估视频大模型多能力的基准,涵盖理解、生成、编辑与重建。

UniVBench: Towards Unified Evaluation for Video Foundation Models

  • 构建跨四大任务的统一评测框架,支持多镜头复杂视频评估。
  • 包含200个真人创作高质量视频,每段配详细描述与多种编辑指令。
  • 开发标准化智能评估系统,确保结果公平可复现,适配研究者与开发者。

视频基础模型旨在将视频理解、生成、编辑和指令跟随整合于单一框架中,是下一代多模态系统的核心方向。然而,现有评测基准仍呈碎片化,仅针对单一任务,依赖特定指标,且多使用短而简单的视频片段,难以体现模型的统一能力。为此,我们提出UniVBench,专为评估视频基础模型在四个核心能力上的表现而设计:视频理解、视频生成、视频编辑,以及新提出的视频重建任务——评估模型对已见过视频内容的还原精度。该基准通过引入200个高质量、多样且多镜头的视频,每个视频配有详尽描述、多格式编辑指令和参考图像,显著提升评估复杂度。所有视频均为人工创作并经严格验证,包含比以往基准更丰富的电影级信息。此外,我们开发了统一的智能评估系统(UniV-Eval),实现提示标准化、指令解析与评分一致,支持对统一视频模型的公平、可扩展、可复现比较。基于指令驱动的多镜头视频任务,UniVBench首次提供了衡量视频基础模型集成能力的完整框架。大量人工标注确保评测结果与人类判断一致,推动稳健视频智能的发展。

原文摘要 · Abstract (English)

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation benchmarks remain fragmented and limited in scope, as they each target a single task, rely on task-specific metrics, and typically use short or simple video clips. As a result, they do not capture the unified capabilities that these models are designed to deliver. To address this gap, we introduce UniVBench, a benchmark purpose-built for evaluating video foundation models across four core abilities: video understanding, video generation, video editing, and a newly proposed task, video reconstruction, which assesses how faithfully a model can reproduce video content it has encountered. Our benchmark substantially expands the complexity of evaluation by incorporating 200 high-quality, diverse and multi-shot videos, each paired with detailed captions, multi-format editing instructions, and reference images. All videos are human-created and carefully validated, offering richer cinematic information than prior benchmarks. In addition, we develop a unified agentic evaluation system (UniV-Eval) that standardizes prompting, instruction parsing, and scoring across all tasks, enabling fair, scalable, and reproducible comparisons of unified video models. By grounding evaluation in instruction-based multi-shot video tasks, UniVBench provides the first framework for measuring the integrated capabilities that video foundation models aim to achieve. Extensive human annotations ensure our evaluation aligns with human judgment, enabling rigorous assessment and accelerating progress toward robust video intelligence.

视频评估多模态基础模型统一评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。