MORSE-500用可编程视频测试多模态推理,突破传统静态数据局限。
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
- 用脚本生成500段可控视频,精确调节复杂度与难度
- 现有多模态模型在抽象和规划任务上差距显著
- 支持持续生成新样本,适合下一代模型压力测试
尽管视觉语言模型(VLMs)快速发展,现有多模态推理基准在三个关键维度存在不足:其一,过度依赖静态图像,无法捕捉现实环境的时间动态;其二,仅聚焦数学求解,忽视抽象、物理、规划、空间与时间等多元推理能力;其三,多数基准很快饱和,难以诊断失败模式或衡量持续进步。我们提出MORSE-500(多模态推理压力测试环境),一个由500个完整脚本化视频组成的视频基准,涵盖六类互补推理类别,每段视频通过确定性Python脚本(使用Manim、Matplotlib、MoviePy)、生成式视频模型及精选实拍素材构建。该脚本驱动设计可精细控制视觉复杂度、干扰项密度与时间动态,实现难度的系统性提升。不同于易饱和的静态基准,MORSE-500具备可进化能力:其可控生成管道支持生成任意挑战性的新实例,非常适合下一代模型的压力测试。初始实验显示,包括Gemini 2.5 Pro、OpenAI o3及主流开源模型在内的前沿系统在所有类别中均存在显著性能差距,尤其在抽象与规划任务上表现薄弱。我们公开完整数据集、生成脚本与评估工具链,推动透明、可复现且面向未来的多模态推理研究。
原文摘要 · Abstract (English)
Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal complexity of real-world environments. Second, they narrowly focus on mathematical problem-solving, neglecting the broader spectrum of reasoning skills -- including abstract, physical, planning, spatial, and temporal capabilities -- required for robust multimodal intelligence. Third, many benchmarks quickly saturate, offering limited headroom for diagnosing failure modes or measuring continued progress. We introduce MORSE-500 (Multimodal Reasoning Stress-test Environment), a video benchmark composed of 500 fully scripted clips with embedded questions spanning six complementary reasoning categories. Each instance is programmatically generated using deterministic Python scripts (via Manim, Matplotlib, MoviePy), generative video models, and curated real footage. This script-driven design allows fine-grained control over visual complexity, distractor density, and temporal dynamics -- enabling difficulty to be scaled systematically as models improve. Unlike static benchmarks that become obsolete once saturated, MORSE-500 is built to evolve: its controllable generation pipeline supports the creation of arbitrarily challenging new instances, making it ideally suited for stress-testing next-generation models. Initial experiments with state-of-the-art systems -- including various Gemini 2.5 Pro and OpenAI o3 which represent the strongest available at the time, alongside strong open-source models -- reveal substantial performance gaps across all categories, with particularly large deficits in abstract and planning tasks. We release the full dataset, generation scripts, and evaluation harness to support transparent, reproducible, and forward-looking multimodal reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。