用新基准测试大模型设计复杂系统的控制器,还蒸馏出可部署的轻量模型。
Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems

- 构建包含132种配置的复杂系统控制评测基准,覆盖多自由度和多种耦合方式。
- 大模型成功率从50%到94.8%,自由度和控制器类型影响最大,差距超36%。
- 通过推理蒸馏得到1.5亿参数轻量模型,可在机器人上稳定完成目标跟踪。
尽管大语言模型(LLMs)在多个科学领域展现出强大能力,但其在反馈控制器设计方面的应用仍待探索。现有基准主要针对线性单自由度(DoF)系统和大型API托管模型,对复杂控制任务的表现及边缘部署可行性缺乏评估。为此,我们提出复杂动力系统至控制的基准(CoDyControlBench),涵盖132种系统配置,五个评估维度:自由度数量、系统类型、耦合程度、阻尼状态和控制器类型。六种先进大模型(GPT、Gemini、Claude、GLM、DeepSeek、Qwen)进行了三次独立运行评估。GPT表现最佳,设计成功率达94.8%,而Qwen最低为50.0%。在各维度中,自由度和控制器类型导致的平均成功率差异最大,分别为36.3%和17.6%,高于系统类型、耦合程度和阻尼状态。对比GPT与Qwen发现,性能差距主要源于控制设计知识,尤其是增益选择和瞬态限制机制的使用。为实现边缘部署,通过推理蒸馏开发了一个1.5亿参数的专用模型。该模型在CoDyControlBench上优于答案蒸馏和基础模型,1-6自由度下表现稳定,并在气动人工肌肉驱动机械臂的三次物理实验中均成功实现目标追踪。
原文摘要 · Abstract (English)
Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchmarks focus mainly on linear single-Degree-of-Freedom (DoF) systems and large API-hosted models, leaving performance on complex controller-design tasks and feasibility for edge deployment unclear. To address these limitations, we introduce the Complex Dynamics-to-Control Benchmark for Large Language Models (CoDyControlBench), comprising 132 system configurations across five evaluation dimensions: number of DoF, system type, coupling level, damping regime, and controller type. Six state-of-the-art LLMs were evaluated over three independent runs, including three commercial models (GPT, Gemini, and Claude) and three open-source models (GLM, DeepSeek, and Qwen). GPT achieved the highest design success rate at 94.8\%, whereas Qwen showed the lowest rate at 50.0\%. Across the benchmark dimensions, DoF and controller type exhibited the largest model-averaged variations in design success, with success-rate ranges of 36.3\% and 17.6\%, respectively, exceeding those associated with system type, coupling level, and damping regime. Comparison of GPT and Qwen showed that their performance gap arose mainly from the control-design knowledge, particularly gain selection and the use of transient-limiting mechanisms. For edge deployment, a specialized 1.5B-parameter model was developed through reasoning distillation. The reasoning-distilled model outperformed the answer-distilled and base model on CoDyControlBench, maintained stable performance across 1-6 DoFs, and achieved successful traget tracking in all three physical trials on a pneumatic-artificial-muscle-driven robotic arm. These results establish a benchmark baseline and highlight the potential of lightweight, edge-deployable controller-design models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。