arXiv:2512.02622cs.CV2025-12被引 9

评测视频生成模型的规则推理能力,发现顶尖模型仅48.87%准确。

RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence

  • 基于文本/图像生成视频,设计40项规则任务覆盖六类认知规则。
  • 用GPT-4o评分,与人工判断一致率达85%,顶尖模型规则一致性仅48.87%。
  • 适合关注视频生成逻辑推理能力的研究者和开发者。

视频生成技术在时间一致性与视觉质量上取得显著进展,正迈向视觉基础模型的关键一步。现有评估基准多聚焦于视觉感知与理解,如美学、指令遵循和时序连贯性,但对视频生成模型的规则推理能力仍缺乏系统探索。尽管已有研究尝试检验模型是否具备零样本学习能力,却缺少细粒度推理能力分解与完整评估协议。为此,我们提出RULER-Bench,一个从认知规则视角评估视频生成模型推理能力的基准。该基准基于文本到视频与图像到视频两大范式,涵盖40个代表性任务,分属六类规则,共622个高质量标注实例。针对每段生成视频,构建包含四项指标的检查清单,并利用GPT-o3对每个问题打分,与人工评判达到85%的一致性。大量实验表明,当前最先进模型在规则一致性指标上仅达48.87%,凸显下一代视频生成模型在推理能力上的巨大提升空间。我们期望RULER-Bench的洞察能推动推理感知型视频生成的发展,助力视频生成模型向视觉基础智能迈进。

原文摘要 · Abstract (English)

Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision foundation models. To evaluate these video generation models, existing benchmarks primarily focus on factors related to visual perception and understanding, like visual aesthetics, instruction adherence, and temporal coherence. However, the rule-based reasoning capabilities of video generation models remain largely unexplored. Although recent studies have carried out preliminary explorations into whether video models can serve as zero-shot learners, they still lack a fine-grained decomposition of reasoning capabilities and a comprehensive evaluation protocol. To address this gap, we introduce RULER-Bench, a benchmark designed to evaluate the reasoning ability of video generation models from the perspective of cognitive rules. Built upon two fundamental paradigms: text-to-video and image-to-video, RULER-Bench covers 40 representative tasks spanning six rule categories with 622 high-quality annotated instances. For the evaluation of each generated video, we construct a checklist covering four metrics and leverage GPT-o3 to assign scores to each question, achieving 85% alignment with human judgements. Extensive experiments show that the state-of-the-art model achieves only 48.87% on the rule coherence metric, highlighting significant room for improvement in the reasoning capability of next-level video models. We expect that the insight obtained from RULER-Bench will facilitate further development of reasoning-aware video generation, advancing video generation models toward vision foundation intelligence.

视频生成规则推理评估基准基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。