用AI让机器人自动插入柔性电缆,无需人工调路径。
Reinforcement Learning for Robotic Insertion of Flexible Cables in Industrial Settings
- 用视觉语言模型自动生成分割掩码,指导强化学习在仿真中训练。
- 在真实环境部署时零样本即用,无需微调,成功率超90%。
- 通过仿真训练规避物理风险,适合工业自动化场景。
工业环境中将柔性扁平电缆(FFCs)插入插座需亚毫米级精度,但其易变形特性使传统机器人操作依赖人工规划轨迹。虽然强化学习(RL)可避免建模复杂力学,但因电缆不确定性导致训练耗时且直接在真实环境训练存在安全风险。为此,本文提出一种基于基础模型的“真实到仿真”训练方法:完全在仿真中进行强化学习,利用语义分割掩码保留电缆与插座的几何空间信息;采用分割任意模型2(SAM2)并结合视觉语言模型(VLM)自动完成初始提示,实现无须人工干预的分割;实验表明该方法具备零样本迁移能力,可直接部署至真实环境,无需微调。
原文摘要 · Abstract (English)
The industrial insertion of flexible flat cables (FFCs) into receptacles presents a significant challenge owing to the need for submillimeter precision when handling the deformable cables. In manufacturing processes, FFC insertion with robotic manipulators often requires laborious human-guided trajectory generation. While Reinforcement Learning (RL) offers a solution to automate this task without modeling complex properties of FFCs, the nondeterminism caused by the deformability of FFCs requires significant efforts and time on training. Moreover, training directly in a real environment is dangerous as industrial robots move fast and possess no safety measure. We propose an RL algorithm for FFC insertion that leverages a foundation model-based real-to-sim approach to reduce the training time and eliminate the risk of physical damages to robots and surroundings. Training is done entirely in simulation, allowing for random exploration without the risk of physical damages. Sim-to-real transfer is achieved through semantic segmentation masks which leave only those visual features relevant to the insertion tasks such as the geometric and spatial information of the cables and receptacles. To enhance generality, we use a foundation model, Segment Anything Model 2 (SAM2). To eleminate human intervention, we employ a Vision-Language Model (VLM) to automate the initial prompting of SAM2 to find segmentation masks. In the experiments, our method exhibits zero-shot capabilities, which enable direct deployments to real environments without fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。