arXiv:2606.27537cs.CV2026-06被引 6

测试视频模型在物体消失后重现时的动态记忆能力。

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

论文配图:MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
图 1 · 摘自论文原文
  • 设计物体消失再出现的动态环境测试场景。
  • 360段真实与合成视频,评估8个主流模型表现。
  • 揭示当前模型在动态变化下的记忆一致性短板。

视频生成模型致力于模拟动态环境,现有基准多仅评估目标在视野内时的记忆一致性;少数考察遮挡的情况,却局限于静态场景。为此,我们提出MemoBench,一个基于物体“消失-重现”范式的诊断性基准:目标经历物理变化后离开视野,重现在更新状态。我们构建了360段包含合成与真实场景的真值视频,并设计涵盖自动化指标与VQA评估的四维诊断体系。对八种先进模型的评估揭示了其在动态变化环境中保持记忆一致性的关键洞见与挑战。

原文摘要 · Abstract (English)

Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few that force objects out of view evaluate static scenes where nothing changes during occlusion. To bridge this gap, we introduce MemoBench, a diagnostic benchmark built around the disappear-and-reappear paradigm in dynamically changing environments: a target object undergoes a physical process, disappears from view, and must be correctly recovered in its updated state upon reappearance. We curate 360 ground-truth clips spanning synthetic and real-world scenes, and design an evaluation suite combining automated metrics with VQA-based assessment across four diagnostic pillars. Evaluation of eight state-of-the-art models reveals key insights and open challenges regarding memory consistency under the disappear-and-reappear paradigm.

视频生成动态建模记忆一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。