让大模型理解视频中多层次动态变化,提升叙事与模拟能力
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

- 用动态架构学习视频关键帧间的状态转移序列
- 在多专家框架下联合预测变化与模拟状态,提升准确性
- 适合需要精准控制视觉动态的生成任务使用者
多模态大模型难以系统建模视频或图像序列中的时间演化。这类输入要求模型预测或模拟多个层次的动态成分,如视觉序列中的动作及其引发的环境变化。为此,我们提出动态模式引导的世界模型DynaVieW,专用于视觉动态预测与模拟。DynaVieW通过学习交错的状态转移序列实现对视觉动态的深入理解:状态涵盖视频关键帧中的广阔视觉场景,转移则捕捉分层模式下的完整动态成分。DynaVieW在混合专家架构下联合建模转移预测与状态模拟,采用跨专家选择性注意力和模式令牌重加权损失,确保有效且鲁棒的学习。该模型对视觉动态的理解显著提升了其在视觉叙事生成与世界模拟任务中的表现,体现出更强的一致性、可控性和指令遵循能力。
原文摘要 · Abstract (English)
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW's understanding of visual dynamics boosts its downstream performance in visual narrative creation and world simulation, showing improved consistency, controllability, and instruction-following.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。