用仿真追踪评估大模型代码语义,提升仪器控制生成质量
EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines
- 通过执行轨迹对齐,模拟真实系统行为评估代码等价性
- 30多个大模型测试显示,部分顶尖模型接近人类控制水平
- 适合科研自动化、智能实验系统开发者参考
评估大语言模型(LLMs)在仪器控制中的表现,需超越传统静态算法基准,因物理系统行为无法仅靠单元测试完整捕捉。本文提出EnvTrace,一种基于仿真的方法,通过执行轨迹对齐来评估代码的语义等价性。该方法在同步辐射光束线控制逻辑的数字孪生环境中验证,数字孪生本身也支持实验前的实时验证。超过30个大模型通过轨迹对齐生成多维度功能正确性评分,结果显示多个顶级模型在快速生成控制代码方面已接近人类水平。这是迈向更广阔愿景的第一步:大模型与数字孪生协同工作——大模型提供直观控制与代理编排,数字孪生提供安全高保真的环境,为自主具身智能铺路。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) for instrument control requires methods that go beyond standard, stateless algorithmic benchmarks, since the behavior of physical systems cannot be fully captured by unit tests alone. Here we introduce EnvTrace, a simulation-based method that evaluates execution traces to assess semantic code equivalence. EnvTrace is demonstrated with a beamline control-logic digital twin to facilitate the evaluation of instrument control code, with the digital twin itself also enabling the pre-execution validation of live experiments. Over 30 LLMs were evaluated using trace alignment to generate a multi-faceted score for functional correctness across key behavioral dimensions, showing that many top-tier models can approach human-level performance in rapid control-code generation. This is a first step toward a broader vision where LLMs and digital twins work symbiotically: LLMs providing intuitive control and agentic orchestration, and digital twins offering safe and high-fidelity environments, paving the way towards autonomous embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。