arXiv:2507.00268cs.ROcs.AI2025-07被引 1

让智能体在训练时预判执行误差,提升真实环境下的控制鲁棒性。

Control-Optimized Deep Reinforcement Learning for Artificially Intelligent Autonomous Systems

  • 构建双阶段框架:先定目标动作,再选补偿控制信号。
  • 在五种机械仿真环境中验证,显著提升实际执行稳定性。
  • 适合机器人、机电系统等需高精度控制的工程应用。

深度强化学习(DRL)已成为复杂决策的强大工具,但传统方法常假设动作能完美执行,忽视了智能体选择动作与系统实际响应之间的不确定性。在机器人、机电系统和通信网络等真实场景中,由系统动态、硬件限制和延迟引起的执行偏差会显著降低性能。本文提出一种控制优化的DRL框架,显式建模并补偿动作执行不匹配问题。该方法采用两阶段结构:先确定期望动作,再选择合适的控制信号以确保正确执行,并在训练中同时考虑动作偏差与控制器修正。通过将这些因素纳入训练过程,智能体可优化动作以兼顾实际控制信号与预期结果,明确处理执行误差。该方法增强了鲁棒性,使决策在真实不确定性下仍有效。我们在五个重构后的开源机械仿真环境中评估,验证了其对不确定性的强适应能力,为控制导向应用提供了高效实用的解决方案。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL) has become a powerful tool for complex decision-making in machine learning and AI. However, traditional methods often assume perfect action execution, overlooking the uncertainties and deviations between an agent's selected actions and the actual system response. In real-world applications, such as robotics, mechatronics, and communication networks, execution mismatches arising from system dynamics, hardware constraints, and latency can significantly degrade performance. This work advances AI by developing a novel control-optimized DRL framework that explicitly models and compensates for action execution mismatches, a challenge largely overlooked in existing methods. Our approach establishes a structured two-stage process: determining the desired action and selecting the appropriate control signal to ensure proper execution. It trains the agent while accounting for action mismatches and controller corrections. By incorporating these factors into the training process, the AI agent optimizes the desired action with respect to both the actual control signal and the intended outcome, explicitly considering execution errors. This approach enhances robustness, ensuring that decision-making remains effective under real-world uncertainties. Our approach offers a substantial advancement for engineering practice by bridging the gap between idealized learning and real-world implementation. It equips intelligent agents operating in engineering environments with the ability to anticipate and adjust for actuation errors and system disturbances during training. We evaluate the framework in five widely used open-source mechanical simulation environments we restructured and developed to reflect real-world operating conditions, showcasing its robustness against uncertainties and offering a highly practical and efficient solution for control-oriented applications.

强化学习控制优化机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。