LLM在复杂控制器调优中仅适合作为结构初始化工具,数据驱动方法更优。
Numbers Beat Words: A Rigorous On-Premise Benchmark for Coupled MIMO Controller Tuning

- 用提示工程引导的LLM可提供可靠初始结构,但依赖特定示例而非推理能力
- 在四类不同系统上,基于开环数据的虚拟参考反馈调优(VRFT)成功率10/10,性能达J=11.12±0.05
- 对良性系统推荐经典方法,有开环数据时首选VRFT,仅当无直接路径时才用LLM
强耦合多输入多输出(MIMO)过程的控制器调优困难,因分散自动调参忽略回路交互,局部优化又敏感于初始值。本文评估本地部署的大语言模型(LLM)是否能提供有效结构先验,并对比经典方法的必要性。单回路CSTR实验中,继电反馈调优优于LLM;在病态四水箱系统中,原始继电法、简单LLM和平衡起始局部优化均失败,而经过提示引导的LLM找到稳定非对称解域,经局部精炼后10次运行均达到J = 12.0 ± 0.16。然而消融实验表明其可靠性更多源于答案式提示样本,而非对耦合数据的推理。另一数据驱动方法虚拟参考反馈调优(VRFT)仅需一次开环实验,无需LLM,经相同精炼后10/10成功,性能提升至J = 11.12 ± 0.05。尽管VRFT需参考模型时间常数τ,采用确定性中位τ规则即可匹配或超越LLM引导选择,且无额外成本。四个不同结构系统的实验表明,从阶跃响应计算的相对增益阵列可预测结构先验是否有效,优化器起始敏感性提供验证信号。结论清晰:对良构系统用经典方法,有开环数据时优先选VRFT,仅当直接路径不可行时才使用LLM。该基准显示LLM价值在于结构而非数值,核心命题‘数字胜过文字’成立。
原文摘要 · Abstract (English)
Tuning controllers for strongly coupled multi-input multi-output (MIMO) processes is difficult because decentralized auto-tuning ignores loop interaction and local optimization is start-sensitive. We benchmark whether on-premise open-weight large language models (LLMs) provide useful structural priors, while testing classical alternatives that may make them unnecessary. On a single-loop CSTR, relay-feedback tuning outperforms the LLM. On a pathological quadruple-tank, naive relay, naive LLM, and balanced-start local optimization fail, whereas a scaffolded LLM finds a reliable asymmetric basin and, after local refinement, reaches J = 12.0 +/- 0.16 in 10/10 runs. However, an ablation shows that this reliability depends more on an answer-shaped prompt example than on reasoning over coupling data. A direct data-driven alternative, Virtual Reference Feedback Tuning (VRFT), uses one open-loop experiment and no LLM; with the same refinement it succeeds in 10/10 runs and improves the result to J = 11.12 +/- 0.05. Although VRFT requires a reference-model time constant tau, a deterministic median-tau rule matches or exceeds LLM-guided selection at no extra cost. Across four structurally different plants, the relative gain array computed from step tests predicts when a structural prior is worth using; optimizer start-sensitivity provides a confirming second signal. The resulting boundary is clear: use classical tuning on benign plants, prefer VRFT when informative open-loop data are available, and reserve LLMs for structural initialization when direct routes are unavailable. The benchmark shows that the LLM's value is structural rather than numerical, and that on the central case, numbers beat words.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。