用视觉语言模型自动优化机器人擦除策略,提升学习效率与效果
Learning a High-quality Robotic Wiping Policy Using Systematic Reward Analysis and Visual-Language Model Based Curriculum
- 通过分析任务收敛性,提出有界奖励机制解决传统方法难题
- 结合视觉语言模型设计课程化训练流程,实现多场景高效学习
- 适用于复杂曲面、不同摩擦条件的高质擦除任务,无需人工调参
自主机器人擦除在工业制造和医疗消毒等领域具有重要意义。尽管深度强化学习展现出潜力,但常面临重复奖励工程的高需求。本文首先分析高质量擦除任务(兼顾清洁效果与完成速度)的收敛性问题,揭示其难以优化,并提出新的有界奖励公式使其可解。进一步,提出一种基于视觉-语言模型(VLM)的课程学习方法,主动监测学习进展并建议超参数调整。实验表明,该方法可在具有不同曲率、摩擦系数和路径点的表面中学习出理想的擦除策略,而基线方法无法实现。项目演示见:https://sites.google.com/view/highqualitywiping。
原文摘要 · Abstract (English)
Autonomous robotic wiping is an important task in various industries, ranging from industrial manufacturing to sanitization in healthcare. Deep reinforcement learning (Deep RL) has emerged as a promising algorithm, however, it often suffers from a high demand for repetitive reward engineering. Instead of relying on manual tuning, we first analyze the convergence of quality-critical robotic wiping, which requires both high-quality wiping and fast task completion, to show the poor convergence of the problem and propose a new bounded reward formulation to make the problem feasible. Then, we further improve the learning process by proposing a novel visual-language model (VLM) based curriculum, which actively monitors the progress and suggests hyperparameter tuning. We demonstrate that the combined method can find a desirable wiping policy on surfaces with various curvatures, frictions, and waypoints, which cannot be learned with the baseline formulation. The demo of this project can be found at: https://sites.google.com/view/highqualitywiping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。