arXiv:2510.04076cs.ROcs.SY2025-10被引 2

提出八种方法提升数据驱动控制的效率与安全,适配实时系统部署。

From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents

  • 融合模型预测、强化学习等八种技术,构建高效安全控制框架
  • 在机械臂、软体机器人和车辆控制中实现低延迟响应(<10ms)
  • 适合计算资源受限的机器人与自动驾驶场景

现代控制应用,尤其是机器人和车辆运动控制,面临实现精确、快速且安全运动的挑战。为此,已发展出最优控制策略以兼顾安全与高性能。由于真实系统的先验物理模型通常可得,基于模型的控制器被广泛使用。模型预测控制(MPC)是主流方法,能在显式处理安全约束的同时优化性能。然而,复杂系统的精确建模困难,促使数据驱动方法兴起。基于机器学习的MPC利用学习模型减少对人工动力学建模的依赖,强化学习(RL)则可直接从交互数据中学习近似最优策略。数据启用预测控制(DeePC)进一步跳过建模环节,直接从原始输入-输出数据中学习安全策略。近年来,大型语言模型(LLM)代理也崭露头角,能将自然语言指令转化为最优控制问题的形式化表达。尽管如此,数据驱动策略仍存在响应慢、计算量大、内存需求高等显著局限,难以应用于具有快速动态、有限机载算力或严格内存约束的现实系统。为此,已有研究提出降阶建模、函数近似策略学习和凸松弛等技术以降低计算复杂度。本文总结并验证了八种此类方法在实际应用中的有效性,涵盖机械臂、软体机器人及车辆运动控制等场景。

原文摘要 · Abstract (English)

One of the main challenges in modern control applications, particularly in robot and vehicle motion control, is achieving accurate, fast, and safe movement. To address this, optimal control policies have been developed to enforce safety while ensuring high performance. Since basic first-principles models of real systems are often available, model-based controllers are widely used. Model predictive control (MPC) is a leading approach that optimizes performance while explicitly handling safety constraints. However, obtaining accurate models for complex systems is difficult, which motivates data-driven alternatives. ML-based MPC leverages learned models to reduce reliance on hand-crafted dynamics, while reinforcement learning (RL) can learn near-optimal policies directly from interaction data. Data-enabled predictive control (DeePC) goes further by bypassing modeling altogether, directly learning safe policies from raw input-output data. Recently, large language model (LLM) agents have also emerged, translating natural language instructions into structured formulations of optimal control problems. Despite these advances, data-driven policies face significant limitations. They often suffer from slow response times, high computational demands, and large memory needs, making them less practical for real-world systems with fast dynamics, limited onboard computing, or strict memory constraints. To address this, various technique, such as reduced-order modeling, function-approximated policy learning, and convex relaxations, have been proposed to reduce computational complexity. In this paper, we present eight such approaches and demonstrate their effectiveness across real-world applications, including robotic arms, soft robots, and vehicle motion control.

控制算法强化学习机器人LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。