arXiv:2602.10840cs.LG2026-02被引 1

用物理动画生成挑战大模型代码能力,首次系统评估并提升其表现。

Training and Benchmarking Code Generation for Physics-Inspired Animations

  • 构建52个物理概念的动画代码数据集,含7659条场景。
  • 最强模型仅21.5%准确率,证明该任务难度高。
  • 引入视觉奖励机制,通过视频验证提升模型生成效果。

大型语言模型在数学推理、复杂编程和科学问题求解等领域已得到广泛研究,但其生成可执行代码以可视化呈现物理情景及其定性动态的能力仍待探索。我们提出SimuScene,这是首个系统性地训练与评估大模型在52个涵盖五个物理领域的物理启发动画代码生成任务上的研究。我们建立了一个自动数据采集管道,并经人工验证确保数据质量。最终数据集包含7,659个场景,其中包含334个经人工验证的测试样本。我们评估了10个主流大模型,发现即使是最强模型也仅达到21.5%的Avg@8准确率,表明生成既可执行又与物理描述视觉一致的动画极具挑战性。最后,我们引入一种基于强化学习的训练流程,利用代码生成视频的视觉反馈作为奖励信号,由视觉-语言模型通过验证问题评估视频质量。实验表明,使用该数据集与视频驱动奖励可显著提升大模型在物理启发动画生成上的性能。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving. However, their ability to generate executable code that visually depicts physical scenarios and their qualitative dynamics remains underexplored. We propose SimuScene, the first systematic study that trains and evaluates LLMs on code generation for physics-inspired animations across 52 concepts spanning five physics domains. We build an automated data collection pipeline with human verification to ensure data quality. The resulting dataset contains 7,659 scenarios, including a 334-example human-verified test set. We evaluate 10 contemporary LLMs and find that even the strongest model achieves only a 21.5\% Avg@8 accuracy, demonstrating the difficulty of generating animations that are both executable and visually aligned with physical scenario descriptions. Finally, we introduce a reinforcement learning pipeline that uses visual rewards from code-generated videos to train text-only LLMs, with a vision-language model evaluating videos through verification questions. Experiments show that training with our data and video-based rewards improves LLM performance on physics-inspired animation generation.

代码生成物理模拟强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。