arXiv:2605.19382cs.AI2026-05被引 1

构建首个大规模程序化视频生成评估基准,揭示代码可运行≠画面空间正确

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

论文配图:PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
图 1 · 摘自论文原文
  • 基于10,372个指令-代码对,覆盖437类主题,支持中英文真实场景
  • 发现主流大模型执行成功率与空间正确率平均相差41%,暴露执行-空间鸿沟
  • 提出四维评估框架,可诊断动态表达与时间密度等关键能力

通过代码生成程序化视频能实现像素级扩散模型无法达到的几何精度与时间连贯性,但如何严格评估语言模型生成的空间正确动画仍是一个开放问题。我们提出PRISM,一个包含10,372个经人工校准的指令-代码对的大规模基准(比之前同类基准大20倍),其场景基于英、中文真实世界知识可视化,覆盖437个主题类别。我们进一步设计了一种漏斗式评估框架,包含四项互补指标:代码级可靠性(可执行性)、空间推理(全程布局正确性)、提示感知动态视觉复杂度(PADVC)和时间密度(TD),用于诊断动态表现力与时间活跃度。对七种主流大模型的系统评估显示显著的执行-空间差距:平均从执行成功率下降至空间通过率约为41%,表明可运行代码并不必然产生空间一致的视觉输出。这些发现表明程序化视频生成评估需超越可执行性。PRISM为推进空间一致的代码生成提供了原则性基准。

原文摘要 · Abstract (English)

Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.

视频生成程序化生成评估基准空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。