用多智能体反馈自动生成物理真实的4D动态场景
PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

- 通过轨迹驱动的多智能体框架,实现力场自动优化
- 生成结果在多样性和物理准确性上显著优于基线
- 适合需要大规模真实物理模拟的科研与创作人群
实现全自动、物理可信的3D运动合成是图形学与生成AI的核心目标。然而,复杂环境力场的配置仍完全依赖人工专家干预,严重制约大规模仿真数据生成。现有自动化方法主要聚焦材料优化,在更复杂的力场优化空间中存在显著模态鸿沟与技术缺陷:朴素大语言模型缺乏模拟反馈,导致严重物理失真;传统评分蒸馏采样(SDS)则面临梯度缓慢、陷入局部最优,且无法动态切换离散力场。为此,我们提出PhysAgent,首个模拟器内嵌的多智能体框架,利用多模态输入实现自动化、物理根基的4D合成。通过解耦内在材质与外在动力学,语义智能体借助外部力场技能模块掌握模拟规则并生成有效初始化。随后,精炼智能体基于轨迹驱动的多智能体反馈,利用视觉基础模型从渲染帧中提取密集点轨迹。将显式运动轨迹转化为结构化文本描述,智能体运用大语言模型常识推理能力执行零样本宏观跃迁,有效逃离局部最优并动态切换离散力场。大量实验表明,PhysAgent能快速从任意多模态提示生成稳定、多样化的物理场景,显著优于现有基线,在生成多样性与物理准确性上均表现卓越。
原文摘要 · Abstract (English)
Achieving fully automated, physically plausible 3D motion synthesis is a core objective in graphics and generative AI. However, configuring complex environmental force fields still relies entirely on manual expert intervention, creating a severe bottleneck for large-scale simulation data generation. Existing automated methods primarily focus on material optimization and exhibit severe modality gaps and technical flaws when applied to the vastly more complex force field optimization space: naive Large Language Models (LLMs) lack underlying simulation feedback, causing severe physical inaccuracies, while traditional Score Distillation Sampling (SDS) suffers from sluggish gradients, local optima entrapment, and a mathematical inability to dynamically switch discrete force fields. To address this, we propose PhysAgent, the first simulator-in-the-loop multi-agent framework that leverages multimodal inputs for automated, physically grounded 4D synthesis. By decoupling intrinsic materials from extrinsic dynamics, PhysAgent utilizes a Semantic Agent equipped with an externalized Force Field Skill module to master simulation rules and generate valid initializations. Subsequently, the Refine Agents, driven by Trajectory-Grounded Multi-Agent Feedback, leverage vision foundation models to extract dense point trajectories from rendered frames. By converting these explicit motion trajectories into structured textual descriptors, the agent harnesses LLM commonsense reasoning to execute zero-shot macroscopic leaps, effectively escaping local optima and dynamically switching discrete force fields. Extensive experiments demonstrate that PhysAgent rapidly generates stable, diverse physical scenes from arbitrary multimodal prompts, significantly outperforming existing baselines in both generation diversity and physical accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。