提出专用于长视频评估的新框架和数据集,解决现有方法对长程连贯性不敏感的问题。
Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

- 设计基于镜头动态的长视频评估指标,捕捉全局叙事一致性。
- 构建包含人类标注的长程特征数据集,实现对长上下文的精准评测。
- 首次将长视频评估作为独立维度,与短视频评估解耦,适合大模型研究者使用。
随着视频生成模型能力不断提升,鲁棒的视频评估指标需求日益迫切。传统指标主要针对短视频设计,侧重帧级视觉质量与局部时序平滑性,难以捕捉长视频的关键特性,如叙事丰富性和全局因果一致性。我们指出短期视觉感知与长上下文属性是本质正交的维度,因此主张将长视频评估从短视频测评中分离。本文提出一套长视频属性破坏测试,揭示现有短视频指标对结构不一致(如镜头扰动、叙事错乱)不敏感的根本缺陷。为此,我们设计基于镜头动态的新度量,对长程测试框架高度敏感。同时引入 Long-CODE(Long-Context as an Orthogonal Dimension for video Evaluation),一个专门用于长视频评估的基准数据集,其人工标注聚焦于真实的长程特征。大量实验表明,所提指标与人类判断达到最先进的相关性。最终,该度量与基准可无缝补充现有短视频标准,建立全面且无偏的视频生成模型评估范式。
原文摘要 · Abstract (English)
As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional metrics are intrinsically tailored for short-video evaluation, predominantly assessing frame-level visual quality and localized temporal smoothness. However, as state-of-the-art video generation models scale to generate longer videos, these metrics fail to capture essential long-range characteristics, such as narrative richness and global causal consistency. Recognizing that short-term visual perception and long-context attributes are fundamentally orthogonal dimensions, we argue that long-video metrics should be disentangled from short-video assessments. In this paper, we focus on the rigorous justification and design of a dedicated framework for long-video evaluation. We first introduce a suite of long-video attribute corruption tests, exposing the critical limitations of existing hort-video metrics from their insensitivity to structural inconsistencies, such as shot-level perturbations and narrative shuffling. To bridge this gap, we design a novel long-video metric based on shot dynamics, which is highly sensitive to the long-range testing framework. Furthermore, we introduce Long-CODE (Long-Context as an Orthogonal Dimension for video Evaluation), a specialized dataset designed to benchmark long-video evaluation, with human annotations isolated specifically to genuine long-range characteristics. Extensive experiments show that our proposed metrics achieve state-of-the-art correlation with human judgments. Ultimately, our metric and benchmark seamlessly complement existing short-video standards, establishing a holistic and unbiased evaluation paradigm for video generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。