arXiv:2510.21758cs.ROcs.LG2025-10综述被引 4

系统梳理强化学习在机器人中的应用与趋势,助力从理论到落地。

Taxonomy and Trends in Reinforcement Learning for Robotics and Control Systems: A Structured Review

  • 构建机器人强化学习应用的结构化分类体系
  • 归纳主流算法如PPO、SAC在连续控制任务中的成效
  • 适合关注机器人智能控制与落地实践的研究者

强化学习(RL)已成为应对动态不确定环境、实现智能机器人行为的基础方法。本文深入回顾了强化学习原理、先进深度强化学习(DRL)算法及其在机器人与控制系统中的集成。从马尔可夫决策过程(MDPs)形式化出发,阐述智能体-环境交互的核心要素,探讨策略梯度、基于价值的学习和演员-评论家方法等核心算法。重点分析DDPG、TD3、PPO、SAC等现代DRL技术在高维连续控制任务中的表现。提出一个结构化分类体系,涵盖运动、操作、多智能体协作及人机交互等应用场景,并对训练方法与部署成熟度进行划分。综述近年研究进展,揭示技术趋势、设计模式及强化学习在真实机器人系统中日益增强的成熟度。旨在连接理论突破与实际应用,为强化学习在自主机器人系统中的演进提供整合视角。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a foundational approach for enabling intelligent robotic behavior in dynamic and uncertain environments. This work presents an in-depth review of RL principles, advanced deep reinforcement learning (DRL) algorithms, and their integration into robotic and control systems. Beginning with the formalism of Markov Decision Processes (MDPs), the study outlines essential elements of the agent-environment interaction and explores core algorithmic strategies including actor-critic methods, value-based learning, and policy gradients. Emphasis is placed on modern DRL techniques such as DDPG, TD3, PPO, and SAC, which have shown promise in solving high-dimensional, continuous control tasks. A structured taxonomy is introduced to categorize RL applications across domains such as locomotion, manipulation, multi-agent coordination, and human-robot interaction, along with training methodologies and deployment readiness levels. The review synthesizes recent research efforts, highlighting technical trends, design patterns, and the growing maturity of RL in real-world robotics. Overall, this work aims to bridge theoretical advances with practical implementations, providing a consolidated perspective on the evolving role of RL in autonomous robotic systems.

强化学习机器人控制综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。