评测视频世界模型的长期记忆能力,发现现有方法存在严重缺陷。
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

- 从实体、环境、因果三方面构建记忆能力评估体系
- 在真实长视频上测试,发现主流模型长期状态保持差
- 适合研究视频生成与世界模型长期一致性的学者
近期基于视频的世界模型在生成高保真视觉序列方面取得突破,但其在长时间跨度下维持稳定合理内部状态的能力仍严重不足。现有基准主要关注视觉质量、运动连贯性和文本-视频对齐,却忽视了世界模型的核心能力——记忆。为此,我们提出 extbf{MBench},一个专注于量化评估视频世界模型记忆能力的综合性基准。我们将记忆能力分解为三个层次互补的核心维度:实体一致性、环境一致性与因果一致性,并进一步细化为12个可量化的子维度,以全面刻画长期记忆表现。该基准基于精心筛选的真实拍摄长视频构建,采用规则化定量指标和视觉语言模型(VLM)进行客观、全面的一致性评估。对主流先进视频世界模型的广泛测试揭示了现有方法在长期状态保持方面的系统性短板,为领域发展提供了标准化评估框架和明确研究方向。
原文摘要 · Abstract (English)
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between visually plausible video generation and the functional requirements of a world model, particularly in maintaining a stable and reasonable internal state over extended temporal horizons. While existing benchmarks primarily emphasize visual quality, motion coherence, and text-video alignment, they largely overlook memory, the core capability of a world model to preserve consistency across long-term horizons and complex interactions. To address this gap, we present \textbf{MBench}, a comprehensive benchmark dedicated to quantifying and evaluating the memory capability of video world models. We systematically decompose the memory capability of video world models into three hierarchical and complementary core dimensions: entity consistency, environment consistency, and causal consistency, which are further refined into 12 quantifiable sub-dimensions for comprehensive characterization of long-term memory. Our benchmark is built upon rigorously curated real-captured long videos, and evaluated by rule-based quantitative matrices and VLM to enable objective and comprehensive consistency assessment. Extensive evaluations of mainstream state-of-the-art video world models reveal critical systemic limitations of existing methods in long-term state retention, providing a standardized benchmark and clear research direction to advance the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。