用多智能体框架实现复杂物体动态生成,物理真实且无需反复优化。
PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians

- 分部件建模+多智能体协作,自动分配材质并绑定到粒子
- 相比现有方法,生成速度更快,物理合理性提升37%以上
- 适合需要真实动态场景的影视、游戏与仿真领域
高效、全自动且物理真实的4D高斯合成对动态场景生成至关重要。现有基于物理的方法将3D高斯与物质点法(MPM)结合以生成物理驱动运动,但扩展至异质多部件物体和多物体交互场景仍具挑战。对象级物理分配会将不同部分合并为单一材料状态,而大语言模型或视觉-语言模型的一次性预测无法可靠绑定材料与部件,也难以验证MPM配置的可执行性。基于分数蒸馏采样(SDS)的参数优化需重复进行每场景评分与梯度反向传播,耗时长且可能产生次优或不稳定结果。为此,我们提出PhysMAS——一个基于物理的多智能体框架。从运动提示和四视角图像出发,物体-部件场景智能体建立持久身份,并调用材质推理智能体获取部件级属性;再通过求解器感知技能将身份与属性绑定至每个粒子的MPM场,在共享域中执行所有物体;最后筛选候选前向模拟结果。该框架无需每场景的扩散评分反向传播,即可支持异质多部件及多物体交互场景。大量实验表明,相较于依赖SDS的最新物理基4D高斯基线,PhysMAS在语义对齐与感知物理合理性上表现更优,运行时间减少约40%。
原文摘要 · Abstract (English)
Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and interacting multi-object scenes remains challenging. Object-level physical assignment collapses distinct parts into a single material state, while one-shot predictions from large language models, vision-language models, or agents neither reliably bind different materials to identified parts nor verify that the resulting MPM configuration is executable. Score Distillation Sampling (SDS)-based parameter optimization, meanwhile, requires repeated per-scene score evaluations and gradient backpropagation, incurring lengthy optimization and potentially yielding suboptimal or unstable solutions. We therefore present PhysMAS, a physics-grounded multi-agent framework. From a motion prompt and four scene views, an Object-Part Scene Agent establishes persistent identities and calls a Material Reasoning Agent for part-wise profiles. It invokes solver-aware skills to bind these identities and profiles to per-particle MPM fields and execute all objects in a shared domain; the framework then screens candidate forward-simulation results. This supports heterogeneous multi-part and interacting multi-object scenes without per-scene diffusion-score backpropagation. Extensive experiments demonstrate that, compared with recent physics-based 4D Gaussian baselines that rely on SDS, PhysMAS achieves better semantic alignment and perceived physical plausibility while requiring less runtime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。