用大模型自动复现控制论文中的仿真,效率提升10倍。
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
- 构建三模块大模型框架,通过迭代反馈提升仿真还原度
- 在500篇论文中成功复现40.7%的仿真结果
- 适合需要快速验证控制算法的研究者
从控制领域论文中重建数值仿真常因参数不全和实现细节模糊而受阻。我们定义了「论文到仿真可复现性」任务,即自动化系统生成能忠实复现论文结果的可执行代码。我们整理了来自IEEE决策与控制会议(CDC)的500篇论文构成基准数据集,并提出RESCORE——一个由Analyzer、Coder和Verifier组成的三组件大模型智能体框架。RESCORE利用迭代执行反馈与视觉对比来提升重建保真度。该方法在基准测试中成功恢复40.7%的合理仿真,优于单次生成方案。值得注意的是,该自动化流程相比人工复现预计实现10倍加速,显著降低验证已发表控制方法的时间与精力成本。我们将公开基准与智能体,以推动研究复现的社区进展。
原文摘要 · Abstract (English)
Reconstructing numerical simulations from control systems research papers is often hindered by underspecified parameters and ambiguous implementation details. We define the task of Paper to Simulation Recoverability, the ability of an automated system to generate executable code that faithfully reproduces a paper's results. We curate a benchmark of 500 papers from the IEEE Conference on Decision and Control (CDC) and propose RESCORE, a three component LLM agentic framework, Analyzer, Coder, and Verifier. RESCORE uses iterative execution feedback and visual comparison to improve reconstruction fidelity. Our method successfully recovers task coherent simulations for 40.7% of benchmark instances, outperforming single pass generation. Notably, the RESCORE automated pipeline achieves an estimated 10X speedup over manual human replication, drastically cutting the time and effort required to verify published control methodologies. We will release our benchmark and agents to foster community progress in automated research replication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。