用自适应混合强化学习实现机器人抛光的精准力控与安全训练
CHEQ-ing the Box: Safe Variable Impedance Learning for Robotic Polishing
- 结合经典控制与强化学习,动态调节机械臂阻抗以适应不同工况
- 硬件实验仅用8小时训练,失败次数控制在5次以内
- 适合需高精度接触作业的工业场景,如打磨、装配
机器人在工业自动化中广泛应用,接触类任务如抛光需要灵巧性和柔顺行为,但传统建模困难。深度强化学习(RL)可通过数据直接学习模型与策略,但存在数据效率低和探索不安全的问题。自适应混合RL方法融合经典控制与强化学习,兼顾结构稳定性与学习能力,提升数据效率与安全性。然而该方法在真实硬件上的应用尚未验证。本文首次在物理机器人上实现了针对变阻抗抛光任务的混合强化学习算法CHEQ。仿真表明,变阻抗可提升抛光性能;对比纯强化学习,CHEQ在保证安全约束下实现有效学习。硬件实验中,仅需8小时训练,发生5次失败即达成有效抛光行为。结果验证了自适应混合强化学习在真实接触任务中的可行性与实用性。
原文摘要 · Abstract (English)
Robotic systems are increasingly employed for industrial automation, with contact-rich tasks like polishing requiring dexterity and compliant behaviour. These tasks are difficult to model, making classical control challenging. Deep reinforcement learning (RL) offers a promising solution by enabling the learning of models and control policies directly from data. However, its application to real-world problems is limited by data inefficiency and unsafe exploration. Adaptive hybrid RL methods blend classical control and RL adaptively, combining the strengths of both: structure from control and learning from RL. This has led to improvements in data efficiency and exploration safety. However, their potential for hardware applications remains underexplored, with no evaluations on physical systems to date. Such evaluations are critical to fully assess the practicality and effectiveness of these methods in real-world settings. This work presents an experimental demonstration of the hybrid RL algorithm CHEQ for robotic polishing with variable impedance, a task requiring precise force and velocity tracking. In simulation, we show that variable impedance enhances polishing performance. We compare standalone RL with adaptive hybrid RL, demonstrating that CHEQ achieves effective learning while adhering to safety constraints. On hardware, CHEQ achieves effective polishing behaviour, requiring only eight hours of training and incurring just five failures. These results highlight the potential of adaptive hybrid RL for real-world, contact-rich tasks trained directly on hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。