arXiv:2602.08025cs.CVcs.AI2026-02被引 20

首个评估世界模型记忆与动作控制能力的闭合回路基准测试

MIND: Benchmarking Memory Consistency and Action Control in World Models

  • 构建250段高分辨率视频,涵盖多视角与动作空间,支持闭环评估
  • 实测显示当前模型难以保持长期记忆一致性且动作泛化能力弱
  • 适合研究视觉记忆、动作规划与世界模型泛化性能的学者使用

世界模型旨在理解、记忆并预测动态视觉环境,但缺乏统一的评估基准。为此,我们提出MIND,首个开放域闭合回路基准,用于评估世界模型在记忆一致性和动作控制方面的能力。MIND包含250段1080p、24 FPS的高质量视频,涵盖100段第一人称+100段第三人称视频(共享动作空间),以及25段+25段跨不同动作空间的视频,覆盖八种多样化场景。我们设计高效评估框架,量化记忆一致性和动作控制能力,捕捉跨视角的时间稳定性与上下文连贯性。此外,通过设计多种动作空间(如不同角色移动速度和相机旋转角度),评估模型在共享场景下对动作空间的泛化能力。为推动后续性能对比,我们引入MIND-World,一种新型交互式视频到世界基线模型。大量实验验证了MIND的完备性,并揭示当前世界模型的关键挑战:长期记忆一致性维持困难,跨动作空间泛化能力不足。代码已开源。

原文摘要 · Abstract (English)

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop revisited benchmark for evaluating Memory consIstency and action coNtrol in worlD models. MIND contains 250 high-quality videos at 1080p and 24 FPS, including 100 (first-person) + 100 (third-person) video clips under a shared action space and 25 + 25 clips across varied action spaces covering eight diverse scenes. We design an efficient evaluation framework to measure two core abilities: memory consistency and action control, capturing temporal stability and contextual coherence across viewpoints. Furthermore, we design various action spaces, including different character movement speeds and camera rotation angles, to evaluate the action generalization capability across different action spaces under shared scenes. To facilitate future performance benchmarking on MIND, we introduce MIND-World, a novel interactive Video-to-World baseline. Extensive experiments demonstrate the completeness of MIND and reveal key challenges in current world models, including the difficulty of maintaining long-term memory consistency and generalizing across action spaces. Code: https://github.com/CSU-JPG/MIND.

世界模型记忆一致性动作控制基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。