arXiv:2510.03727cs.AIcs.CL2025-10被引 1

让多模态大模型具备世界模型能力,能推理、模拟动态、可控生成。

Bridging the Gap Between Multimodal Foundation Models and World Models

论文配图:Bridging the Gap Between Multimodal Foundation Models and World Models
图 1 · 摘自论文原文
  • 通过因果推断等结构化推理提升模型深层理解力。
  • 实现图像视频的可控生成,支持语义一致与用户意图精准匹配。
  • 拓展至4D可控生成,支持时空交互与对象编辑。

人类通过多感官整合理解世界,能够感知、推理并想象动态物理过程。受此启发,多模态基础模型(MFMs)已成为强大的多模态理解和生成工具。然而,当前MFMs难以充当有效世界模型,缺乏反事实推理、动态模拟、时空信息理解、生成结果控制及多维度推理等核心能力。本文探究如何弥合多模态基础模型与世界模型之间的差距。首先,通过判别性任务提升推理能力,并赋予模型因果推断、反事实思维和时空推理等结构化推理技能,使其超越表面相关性,深入理解视觉与文本数据间的内在关联。其次,探索图像与视频模态下的生成能力,提出新的结构化、可控生成框架。方法融合场景图、多模态条件与对齐策略,引导生成过程,确保与高层语义及细粒度用户意图保持一致。进一步将技术扩展至可控4D生成,实现时空上可交互、可编辑、可变形的对象合成。

原文摘要 · Abstract (English)

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation models (MFMs) have emerged as powerful tools for multimodal understanding and generation. However, today's MFMs fall short of serving as effective world models. They lack the essential ability such as perform counterfactual reasoning, simulate dynamics, understand the spatiotemporal information, control generated visual outcomes, and perform multifaceted reasoning. We investigates what it takes to bridge the gap between multimodal foundation models and world models. We begin by improving the reasoning capabilities of MFMs through discriminative tasks and equipping MFMs with structured reasoning skills, such as causal inference, counterfactual thinking, and spatiotemporal reasoning, enabling them to go beyond surface correlations and understand deeper relationships within visual and textual data. Next, we explore generative capabilities of multimodal foundation models across both image and video modalities, introducing new frameworks for structured and controllable generation. Our approaches incorporate scene graphs, multimodal conditioning, and multimodal alignment strategies to guide the generation process, ensuring consistency with high-level semantics and fine-grained user intent. We further extend these techniques to controllable 4D generation, enabling interactive, editable, and morphable object synthesis over time and space.

多模态世界模型可控生成4D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。