用AI自动生成可交互的4D物理世界,让语言描述变真实场景。
GS-Agent: Creating 4D Physical Worlds With Generative Simulation

- 分角色协作的智能体系统,结合物理引擎实现自动化构建。
- 能生成液体、柔体与刚体间的丰富动态交互,支持电影级镜头控制。
- 适合影视创作、游戏设计及物理仿真研究者使用。
从自然语言描述生成动态且物理真实的4D世界既有趣又具挑战性。传统计算机图形学依赖人工制作,需大量人力调整材质、运动和视觉质量。近年生成式基础模型兴起,尝试从大规模数据中学习生成4D世界,但现有方法仍难以保证物理合理性与可控性。本文提出GS-Agent,一个端到端的多智能体框架,通过在环集成物理引擎,将自然语言转化为真实、动态且可控的4D物理世界。受人类构建4D世界方式启发,GS-Agent将任务分解为实体管理(包括3D资产筛选、材质调节、位置摆放与运动控制)和渲染配置(如相机与光照调整)。多个具备不同专长的智能体通过代码与物理引擎交互,获取多模态反馈并协同迭代,逐步构建符合描述的4D世界。实验表明,该系统能有效生成具有丰富液体、柔体与刚体交互的多样化、物理合理的4D世界,并实现电影级相机与灯光控制。我们期望GS-Agent成为4D世界生成的新范式,推动创意内容生成与物理人工智能发展。项目页面:https://umass-embodied-agi.github.io/gs-agent/
原文摘要 · Abstract (English)
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic system that emulates how humans traditionally create 4D worlds, yet automates the entire process. We present GS-Agent, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language. Inspired by how humans build 4D worlds, GS-Agent decomposes the task into entity management, covering 3D asset curation, material tuning, placement, and motion control, and rendering configuration, including camera and lighting manipulation. Multiple agents with distinct expertise interact with the physics engine via code, seek multimodal feedback, and collaborate to iteratively construct 4D worlds that align with the given descriptions. Experimental results show that GS-Agent effectively converts natural language into diverse and physically plausible 4D worlds exhibiting rich interactions among liquids, deformable objects, and rigid bodies, while achieving cinematic camera and lighting control. We envision GS-Agent as a foundation for a new paradigm in 4D world generation, empowering creative content creation and physical AI. Project page at https://umass-embodied-agi.github.io/gs-agent/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。