arXiv:2604.23580cs.ROcs.AI2026-04被引 4

用多智能体迭代纠错提升物理仿真代码生成准确率

PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement

论文配图:PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement
图 1 · 摘自论文原文
  • 三智能体协作:生成、纠错、优化,分步完善仿真代码
  • 新基准测试中达67.7分,比最佳基线高31.4分
  • 适合机器人、具身AI研究者,解决物理描述到代码的鸿沟

物理感知符号化3D场景模拟对机器人、具身AI和科学计算至关重要,要求模型理解自然语言中的物理现象并转化为可执行仿真环境。尽管大语言模型(LLMs)在通用代码生成上表现优异,但在物理描述与仿真实现之间存在语义鸿沟。我们提出PhysCodeBench,首个全面评估物理感知符号化模拟的基准,包含700个手工构建的多样化样本,涵盖力学、流体动力学和软体物理,并配有专家标注。评估框架通过自动化与视觉评估衡量代码可执行性与物理准确性。基于此,我们提出自校正多智能体精炼框架(SMRF),包含三个专用智能体:仿真生成器、错误纠正器和仿真优化器,通过领域特定验证进行迭代协作。SMRF在整体性能上达到67.7分,优于最优基线模型的36.3分,提升31.4分。分析表明,错误纠正对物理准确模拟至关重要,且专用多智能体方法在所测物理领域显著优于单智能体方法。

原文摘要 · Abstract (English)

Physics-aware symbolic simulation of 3D scenes is critical for robotics, embodied AI, and scientific computing, requiring models to understand natural language descriptions of physical phenomena and translate them into executable simulation environments. While large language models (LLMs) excel at general code generation, they struggle with the semantic gap between physical descriptions and simulation implementation. We introduce PhysCodeBench, the first comprehensive benchmark for evaluating physics-aware symbolic simulation, comprising 700 manually-crafted diverse samples across mechanics, fluid dynamics, and soft-body physics with expert annotations. Our evaluation framework measures both code executability and physical accuracy through automated and visual assessment. Building on this, we propose a Self-Corrective Multi-Agent Refinement Framework (SMRF) with three specialized agents (simulation generator, error corrector, and simulation refiner) that collaborate iteratively with domain-specific validation to produce physically accurate simulations. SMRF achieves 67.7 points overall performance compared to 36.3 points for the best baseline among evaluated SOTA models, representing a 31.4-point improvement. Our analysis demonstrates that error correction is critical for accurate physics-aware symbolic simulation and that specialized multi-agent approaches significantly outperform single-agent methods across the tested physical domains.

物理仿真多智能体代码生成具身AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。