视频模型开始具备解棋盘、迷宫等推理能力,成功率超六成。
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
- 用任务对设计实验范式,评估视频模型推理能力。
- Sora-2在五类任务中达60%正确率,验证模型推理潜力。
- 开源评估框架支持高效扩展,适合研究视频模型推理。
我们证明视频生成模型现已具备推理能力。在国际象棋、迷宫、数独、心理旋转和瑞文矩阵等任务上,如Sora-2等领先模型已实现60%的成功率。我们建立了一套以“任务对”为核心的设计范式,并构建了包含39个模型的代码框架,支持高效扩展——用户可便捷添加新模型与任务。实验表明,自动化评估结果与人类判断高度相关,该范式具有强可扩展性。鉴于此,我们提出利用该范式开展强化学习以进一步提升视频模型的推理能力。完整原始结果及开源代码库(VMEvalKit)可访问:https://grow-ai-like-a-child.com/video-reason/ 与 https://github.com/hokindeng/VMEvalKit。
原文摘要 · Abstract (English)
We show that video generation models could reason now. Testing on tasks such as chess, maze, Sudoku, mental rotation, and Raven's Matrices, leading models such as Sora-2 achieve sixty percent success rates. We establish a robust experimental paradigm centered on the "Task Pair" design. We build a code framework, with 39 models available already, that supports this paradigm and allows for easy scaling - users can add models and tasks efficiently. We show our automated evaluation strongly correlates with human judgment, and therefore this paradigm is highly scalable. We see an opportunity, given the availability of our paradigm, to do reinforcement learning for improving reasoning in video models. You could checkout all of our raw $\href{https://grow-ai-like-a-child.com/video-reason/}{results}$ and our $\href{https://github.com/hokindeng/VMEvalKit}{VMEvalKit}$ codebase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。