用自然语言生成可交互编辑的4D动态场景,支持多视角和物体控制。
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
- 通过语言指令生成具多视角一致性的4D动态场景。
- 支持物体定向、变色、删除等实时编辑,无需重生成。
- 适合机器人训练与任务设计,提升数据可控性与复现性。
支持可控且可编辑时空环境的世界模型对机器人技术至关重要,可实现可扩展的训练数据、可复现的评估以及灵活的任务设计。尽管近期文本到视频模型能生成逼真的动态效果,但受限于2D视角且交互能力有限。我们提出MorphoSim,一种基于语言引导的框架,能够生成具备多视角一致性与物体级控制的4D场景。用户可通过自然语言指令,操控场景中物体的运动、颜色或移除,并从任意视角观察。该框架结合轨迹引导生成与特征场蒸馏,实现交互式编辑而无需全量重生成。实验表明,MorphoSim在保持高场景保真度的同时,显著提升了可控性与可编辑性。代码已开源:https://github.com/eric-ai-lab/Morph4D。
原文摘要 · Abstract (English)
World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video models generate realistic dynam ics, they are constrained to 2D views and offer limited interaction. We introduce MorphoSim, a language guided framework that generates 4D scenes with multi-view consistency and object-level controls. From natural language instructions, MorphoSim produces dynamic environments where objects can be directed, recolored, or removed, and scenes can be observed from arbitrary viewpoints. The framework integrates trajectory-guided generation with feature field dis tillation, allowing edits to be applied interactively without full re-generation. Experiments show that Mor phoSim maintains high scene fidelity while enabling controllability and editability. The code is available at https://github.com/eric-ai-lab/Morph4D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。